AI News August 3, 2026: Microsoft Project Perception Launches, OpenAI Astra Solves 10 Open Math Problems, GPT-5.6 Price Cuts, DeepSeek V4 Flash, and AI Compute Doubles
Microsoft's Project Perception enters public preview today with the MAI-Cyber-1-Flash model beating every frontier lab on the CyberGym benchmark at half the cost. OpenAI's unreleased Astra model solved ten previously-open problems in mathematics and theoretical computer science — verified by machine-checkable Lean proofs — for roughly $2,000 in compute. GPT-5.6 Luna dropped 80% in price. DeepSeek's V4 Flash proved frontier agents cost $0.14/M tokens. And the New York Times reports AI chip deployments are now doubling every nine months.
🛡️ Top 5 AI Stories — August 3, 2026
Today’s AI news is defined by three converging forces: specialized models beating general-purpose ones at specific tasks, the collapse of price per unit of intelligence, and AI crossing from benchmark-chasing into genuine original research. Microsoft’s Project Perception enters public preview today, proving that a purpose-built cybersecurity model can outperform every frontier generalist on the most credible security benchmark. OpenAI’s Astra model quietly solved ten open problems in pure mathematics — verified by Lean proofs that any mathematician can independently check. GPT-5.6 and DeepSeek V4 Flash continued the relentless price compression. And the physical substrate of it all — AI chips — is now doubling every nine months. Here’s what matters.
1. 🛡️ Microsoft Project Perception Launches — MAI-Cyber-1-Flash Enters Public Preview Today
On August 3, 2026, Microsoft’s Project Perception enters public preview, alongside MAI-Cyber-1-Flash — the company’s first cybersecurity-specialized AI model. The launch is Microsoft’s most aggressive move yet to capture the AI-driven security market under CEO AI Mustafa Suleyman and newly returned security-unit head Hayete Gallot.
Project Perception is an agentic security platform that coordinates three classes of autonomous AI agents in a continuous loop: red agents that find attack paths and simulate offensive techniques, blue agents that evaluate and prioritize risk, and green agents that generate and deploy patches. The goal is to compress hours of manual security work into machine-speed cycles.
The numbers are striking. On the CyberGym benchmark — which tests whether AI agents can reproduce real software vulnerabilities from source code — MAI-Cyber-1-Flash scored 95.95%. The next-best result, GPT-5.5 Cyber, scored 85.6%. Anthropic’s Mythos, Gemini 3.5, and GPT-5.6 Sol all cluster around 83–84%. Microsoft claims this performance at roughly 50% of the operating cost of the previous MDASH configuration.
A key architectural detail: MAI-Cyber-1-Flash sits inside MDASH, Microsoft’s multi-model vulnerability harness coordinating over 100 specialized agents. The model handles about 90–95% of routine vulnerability queries, escalating the hardest 10% to GPT-5.4. Critically, the harness, security context, and action space are kept separate from the model, making the underlying model swappable without rebuilding the workflow. Access will be through Azure AI Foundry with Microsoft’s existing customer-vetting processes. Human oversight remains central to the design.
Why it matters: This is the clearest evidence yet that specialized models beat general-purpose frontier models on domain-specific tasks — not by a rounding error, but by 12 benchmark points at half the cost. It signals a strategic shift from “one model to rule them all” to vertically integrated, domain-optimized AI stacks.
2. 🔢 OpenAI Astra Solves 10 Open Math Problems — Verified by Lean Proofs for $2,000
On August 1, 2026, OpenAI announced that an internal version of Astra — its next major model family — solved ten previously-open problems across mathematics and theoretical computer science. This is how OpenAI chose to introduce Astra to the world: not with benchmark scores, but with genuine mathematical discovery.
The results span high-dimensional geometry, coding theory, arithmetic circuit complexity, group theory, operator algebras, quantum complexity, lattice cryptography, and extremal combinatorics. Highlights include:
- Non-sofic groups — a construction proving their existence, resolving a central open question in group theory
- Connes’s rigidity conjecture — disproved
- Sphere-packing bounds — new upper bounds down to the Cohn–Elkies threshold
- Multicolor Ramsey numbers — a superexponential lower bound resolving Erdős problem 183
- Closest vector problem — polynomial-factor hardness of approximation, relevant to post-quantum cryptography
What makes this credible rather than hype: OpenAI published machine-checkable Lean proofs on GitHub, so any mathematician can independently verify every logical step. Fields Medalist Timothy Gowers said he would recommend one of the proofs for publication in the Annals of Mathematics “without hesitation.” And the total compute cost was roughly $2,000 at Sol API rates — reframing advanced mathematical research as something scalable with compute rather than gated by scarce human genius.
Why it matters: This is AI crossing from doing tasks into doing original research in the most rigorous field there is — and doing it verifiably. The Lean proofs transform an extraordinary claim into a checkable fact. If this generalizes, the bottleneck on certain kinds of mathematical progress shifts from human genius to available compute.
3. 💰 GPT-5.6 Price War — Luna Down 80%, Terra Down 20%
On July 30, 2026, OpenAI cut GPT-5.6 Luna prices by 80% and GPT-5.6 Terra by 20%, intensifying the frontier model price war. New API pricing: Terra at $2/M input and $12/M output; Luna at $0.20/M input and $1.20/M output. Sol pricing remains unchanged.
The GPT-5.6 family — released July 9, 2026 after a limited preview — has three tiers: Sol (flagship, $5/$30 per M tokens), Terra (balanced, now $2/$12), and Luna (high-volume, now $0.20/$1.20). The lineup was initially held back by a U.S. government request before broader rollout. According to Artificial Analysis, even the Luna model now outperforms Gemini 3.6 Flash, making OpenAI’s cost-per-intelligence increasingly favorable.
GPT-5.6 also introduced OpenAI’s first cache-write pricing (1.25× input rate) while retaining the 90% discount on cache reads. ChatGPT and Codex subscription prices and quota budgets remain unchanged, though Terra and Luna now consume fewer credits.
Why it matters: Frontier intelligence is compressing in price faster than it’s improving in raw capability. When the cheapest frontier-tier model undercuts most competitors and still beats them on benchmarks, the unit economics of building on AI shift decisively toward developers.
4. 🚀 DeepSeek V4 Flash — Frontier Agents at $0.14/M Tokens
On July 31, DeepSeek shipped the official V4 Flash API into public beta — a re-post-trained model that posts massive gains on agent benchmarks while retaining its 284B MoE / 13B-active architecture. The model ID deepseek-v4-flash now natively supports the Responses API format and is adapted for Codex.
On the Artificial Analysis Intelligence Index v4.1, V4 Flash 0731 scored 50 — a 10-point jump over the previous V4 Flash and 6 points ahead of DeepSeek’s own V4-Pro. On Terminal Bench 2.0, it hit 56.9%; on GPQA Diamond, 88.1%. Pricing holds at $0.14/M input and $0.28/M output tokens — roughly 1/20th the cost of frontier closed models — with cached input at just $0.003/M. At 117 tokens/second, it’s also fast.
This reinforces 2026’s defining trend: the frontier is commoditizing from below. A Chinese open-weight model, retrained only at the post-training stage, matches or beats multi-billion-dollar frontier systems on agentic tasks at budget pricing.
5. 📈 AI Compute Doubles Every Nine Months — NYT/Epoch Analysis
The New York Times reported on August 1 that AI chip deployments are now projected to double every nine months, according to Epoch AI data. This cadence outpaces the classic two-year Moore’s Law cycle and is reshaping the global compute landscape — driving massive capital expenditure from hyperscalers, straining energy grids, and intensifying the geopolitical race for advanced semiconductors.
The implication for the AI arms race is direct: model capability is increasingly a function of deployable compute, and the players who can secure the most chips fastest gain a compounding edge. It also explains the simultaneous pressure on pricing (more compute → cheaper inference) and on governance (more capability → harder to regulate).
📌 The Takeaway
Three patterns define August 3, 2026. First, specialization wins: Microsoft’s purpose-built cyber model didn’t edge out generalists — it beat them by double digits at half the cost. The era of one frontier model for everything is giving way to vertically integrated, domain-optimized stacks. Second, price collapse is accelerating: GPT-5.6 Luna at $0.20/$1.20 and DeepSeek V4 Flash at $0.14/$0.28 mean frontier-grade intelligence is now priced like infrastructure, not premium software. Third, AI is doing original research: OpenAI’s Astra didn’t score well on a test — it produced new mathematical knowledge, verified by proofs anyone can check. Combined with compute doubling every nine months, the trajectory is unmistakable. The question is no longer whether AI can contribute to frontier research and specialized professional work. It’s how fast the rest of the economy adapts.
Sources: Microsoft / GeniusFirms, OpenAI, OpenAI (price-performance), Artificial Analysis, New York Times. Read time ~7 min. For daily AI coverage, bookmark this page.
📡 Sources
- ▸ Microsoft — Project Perception & MAI-Cyber-1-Flash announcement (public preview Aug 3, 2026)
- ▸ OpenAI — Ten advances in mathematics and theoretical computer science
- ▸ OpenAI — Advancing the price-performance frontier with GPT-5.6 (Jul 30, 2026)
- ▸ Artificial Analysis — DeepSeek V4 Flash 0731 scores 50 on the Intelligence Index
- ▸ New York Times — AI chip deployments projected to double every nine months