Sol really is a notch up on the hard reasoning and coding benchmarks, but the scoreboards are noisy enough that ‘early AGI’ is mostly a vibes read, not a measurement. Chinese and open-ish stacks (DeepSeek, GLM, Qwen, Hy3, Nemotron) are quietly taking over the cost/perf frontier while big labs and governments turn model access and logging into geopolitical and regulatory battlegrounds.
The net result is a world where raw capability is racing ahead, but what normal people see—Copilot in Excel, AI in Office—is still throttled by trust, tokens, and infrastructure politics.
Key Events
/OpenAI's GPT‑5.6 Sol set a new state‑of‑the‑art on ARC‑AGI‑3 with a 7.8% score.
/SpaceXAI’s Grok 4.5 launched as an Opus‑class coding and agent model, taking #1 on AutomationBench‑AA at low per‑task cost.
/Tencent released Hy3, a 295B‑parameter Apache‑2.0 model with 256k context and a halved hallucination rate for agentic workflows.
/NVIDIA’s Nemotron family passed 100M downloads and underpins 145 ICML papers, with Nemotron 3 Ultra delivering benchmark‑leading performance at roughly 10× lower inference cost than leading proprietary models.
/Anthropic extended Claude Fable 5 access while accusing Alibaba of large‑scale model distillation, as Alibaba banned Claude Code over alleged backdoors.
Report
The frontier story this week isn’t that Sol looks a bit AGI‑ish; it’s that the benchmarks making it look that way crash into very ordinary experiences with coding agents and office AI.
Under the SOTA headlines you can see three races—reasoning, coding, and infra sovereignty—pulling apart from the neat “one best model” narrative.
sol, fable, grok: frontier is moving, but inside a benchmark-shaped box
GPT‑5.6 Sol posts 92.5% on ARC‑AGI‑2 and 91.9% on TerminalBench, and sets a 7.8% SOTA on ARC‑AGI‑3, clearly ahead of GPT‑5.5 on paper.
On DeepSWE, Sol hits 76%, while Grok 4.5 lands at 62%, putting both above most of the pack for coding‑like reasoning. Claude Fable 5 scores 59 on one Intelligence Index and leads EnterpriseOps‑Gym‑AA, but its Pass@1 on a coding benchmark slid from 65.5% in June to 54.8% in July.
Absolute scores on ARC‑AGI‑3 are still single‑digits and SWE‑Bench Pro just admitted ~30% of its tasks are broken, so the gap between “dominates this benchmark” and “dominates reality” is very much intact.
coding and agents: arms race on top of broken scoreboards
Sol via Codex now tops the Artificial Analysis Coding Agent Index with 80 points, while Fable 5 sits just below it and Grok 4.5 wins AutomationBench‑AA on real workflow completion.
Grok 4.5 is reported to match GPT‑5.5‑xhigh on coding at roughly half the cost, and it leads Harvey’s legal agent benchmark, staking out “serious work” territory.
Meta’s Muse Spark 1.1, Tencent’s Hy3, and NVIDIA’s Nemotron‑Puzzle are all marketed as agent‑ and coding‑centric models with million‑token or quarter‑million contexts, but early user reports question whether they actually beat Sol/Fable/Grok outside curated evals.
SWE‑Bench Pro retracting its own recommendation and CursorBench advantaging Grok by training on an older Cursor snapshot are reminders that a lot of the leaderboard spread here is as much artifact as capability.
open(-ish) blocs, chinese cost fronts, and stack fracture
Chinese and quasi‑open stacks—DeepSeek V4, GLM 5.2, Qwen 3.6, Kimi K2.7—now define the cost/perf frontier on many routers, with Chinese models taking over 45% of OpenRouter token volume while OpenAI sits at 7.4%.
GLM 5.2 is reported ~5× cheaper than Opus 4.8 and 11× cheaper than Fable 5 on some workloads while still topping PostTrainBench, and DeepSeek V4 Flash matches Opus‑class coding at lower cost.
At the same time Tencent’s Hy3 (Apache‑2.0), NVIDIA’s Nemotron (100M+ downloads, 10× cheaper Ultra), and Palantir+Nemotron deployments in classified systems are building open‑tunable blocs that don’t depend on OpenAI or Anthropic at all.
Against that, Anthropic is accusing Alibaba of running the largest known distillation campaign on Claude while Alibaba bans Claude Code as a backdoor and China openly discusses restricting overseas access to top models, turning model access itself into a geopolitical lever.
agents, infra, and the token black hole
LangChain’s OpenWiki, Deep Agents, and LangSmith, plus LangGraph in Tier‑1 banks, MCP servers, and registries like Agentshive, suggest agent frameworks are finally past the toy phase.
But the loudest data points are about costs and opacity: a LangGraph agent looping on a broken tool generated a surprise cloud bill, while one LiteLLM user saw their LLM bill triple over eight months with no clear attribution.
On the infra side, vLLM pushes 668 tok/s on Qwen 3.6‑27B FP8 and about 2000 TPS when batching, while llama.cpp’s DFlash and NVFP4 unlock big speedups on local hardware.
Yet GPU utilization across the industry is only 5–10%, DGX Spark owners complain about memory‑bandwidth bottlenecks and $4k‑ish entry costs, and Meta is preparing to rent out its excess compute via Meta Compute.
The net effect is that tokens, not clever agent logic, are what break budgets or make them viable.
safety, misuse, and the adoption paradox
OpenAI is under threat of sanctions over allegedly hiding and deleting ChatGPT logs in the New York Times case, at the same time EU ‘Chat Control’ pushes for scanning private messages and Sentinel‑style gateways separate data from instructions to fight prompt injection.
Claude Mythos Preview coincided with a spike to ~1,500 high‑ and critical‑severity CVEs disclosed in June, GitHub’s AI agent leaked private repos, and a 15‑year‑old was arrested for cyberattacks using ChatGPT‑generated malware.
Alibaba banning Claude Code as a supposed backdoor and the rise of uncensored/local models for privacy show how quickly “safety” arguments flip into “we don’t trust your data path” arguments.
Meanwhile Microsoft 365 Copilot sits under 4.5% adoption with only 1% weekly active, Excel users complain that AI can’t handle complex tasks, MIT data shows weaker brain connectivity when students use ChatGPT to write, and executives are shocked by high AI bills that don’t replace headcount.
Frontier models are racing ahead on math, code, and agents, but the places where people actually live—email, docs, spreadsheets—are moving on a much slower, more suspicious timeline.
What This Means
The visible plot is Sol‑vs‑Fable‑vs‑Grok, but the deeper story is a three‑way divergence between benchmark‑driven capability gains, fragmenting model/compute blocs, and messy deployment economics where tokens, trust, and governance matter more than raw IQ points.
On Watch
/Meta’s Muse Spark 1.1 plus Meta Compute—cheap 1M‑context agentic models bundled with surplus infra—could turn Meta from a model also‑ran into a major low‑end cloud AI vendor if real‑world performance catches up to the marketing.
/China’s discussion of restricting overseas access to its top models, alongside already dominant token share from DeepSeek/GLM/Qwen on routers, is an early sign that model availability itself could become a regulated export good.
/Runaway‑cost stories around LangGraph and LiteLLM, plus new local observability tools and LangGraph Sync, hint that the next wave of agent failures will be about governance and tracing rather than clever orchestration logic.
Interesting
/DeepSeek v4 (Flash) has 284B parameters, making it cheaper to run than smaller models like the 27B Qwen.
/Gemini 3.5 Flash leads in VQA/OCR benchmarks with 90.6% accuracy, outperforming Claude Fable 5 and GPT-5.5.
/GLM 5.2 generated most of a playable 3D game in its first iteration, showcasing its coding capabilities.
/Only 18 cents of each dollar spent on AI coding reaches users as a shipped product due to the costs of fixing coding mistakes.
/LLMs are found to predict other models' outputs better than actual truths, a topic discussed in a paper at ICML.
We processed 10,000+ comments and posts to generate this report.
AI-generated content. Verify critical information independently.
/OpenAI's GPT‑5.6 Sol set a new state‑of‑the‑art on ARC‑AGI‑3 with a 7.8% score.
/SpaceXAI’s Grok 4.5 launched as an Opus‑class coding and agent model, taking #1 on AutomationBench‑AA at low per‑task cost.
/Tencent released Hy3, a 295B‑parameter Apache‑2.0 model with 256k context and a halved hallucination rate for agentic workflows.
/NVIDIA’s Nemotron family passed 100M downloads and underpins 145 ICML papers, with Nemotron 3 Ultra delivering benchmark‑leading performance at roughly 10× lower inference cost than leading proprietary models.
/Anthropic extended Claude Fable 5 access while accusing Alibaba of large‑scale model distillation, as Alibaba banned Claude Code over alleged backdoors.
On Watch
/Meta’s Muse Spark 1.1 plus Meta Compute—cheap 1M‑context agentic models bundled with surplus infra—could turn Meta from a model also‑ran into a major low‑end cloud AI vendor if real‑world performance catches up to the marketing.
/China’s discussion of restricting overseas access to its top models, alongside already dominant token share from DeepSeek/GLM/Qwen on routers, is an early sign that model availability itself could become a regulated export good.
/Runaway‑cost stories around LangGraph and LiteLLM, plus new local observability tools and LangGraph Sync, hint that the next wave of agent failures will be about governance and tracing rather than clever orchestration logic.
Interesting
/DeepSeek v4 (Flash) has 284B parameters, making it cheaper to run than smaller models like the 27B Qwen.
/Gemini 3.5 Flash leads in VQA/OCR benchmarks with 90.6% accuracy, outperforming Claude Fable 5 and GPT-5.5.
/GLM 5.2 generated most of a playable 3D game in its first iteration, showcasing its coding capabilities.
/Only 18 cents of each dollar spent on AI coding reaches users as a shipped product due to the costs of fixing coding mistakes.
/LLMs are found to predict other models' outputs better than actual truths, a topic discussed in a paper at ICML.