Frontier models like Fable 5 are getting more powerful but also more politically locked down, while open and local models are starting to match them on real work. At the same time, token economics are cracking, so the future looks less like one omnipotent AGI chatbot and more like a messy mix of regulated frontier APIs, cheap open models, and hyper-optimized local stacks.
Underneath the AGI discourse, the real story is about who controls access, cost, and liability.
Key Events
/Anthropic launched Claude Fable 5 as its first Mythos-class model, then had Fable 5 and Mythos 5 globally disabled under a U.S. export-control directive within about 72 hours.
/DeepSeek V4 Pro, a 1.6T-parameter open model, surpassed GPT‑5.5 Pro on precision and topped coding benchmarks like SWE-Bench and LiveCodeBench.
/Google DeepMind released DiffusionGemma, a diffusion-based text model that generates blocks of text up to 4× faster, exceeding 1,000 tokens per second on a single H100.
/Apple adopted a new AI architecture pairing on-device CoreAI/MLX models with cloud intelligence based on Google Gemini, including a revamped Siri.
/Kimi K2.7‑Code launched as an open-source coding model with +21.8% to +31.5% benchmark gains over K2.6 and significantly better token efficiency.
Report
Everyone is arguing about whether Claude Fable 5 is 'AGI', but the more interesting story is that frontier models just became politically contingent infrastructure rather than products.
At the same time, open and local stacks are quietly catching up on capability while the economics of tokenmaxxing start to fall apart.
the mythos-class moment and the end of unconditional frontier access
Claude Fable 5 arrived as Anthropic’s first Mythos-class model, posting a 91/100 Senior Engineer score and ranking #1 on GDPval-AA, which made it look like the next step-change in capability.
Within roughly 72 hours, the U.S. government ordered Fable 5 and Mythos 5 globally disabled and imposed export controls blocking foreign access to Anthropic’s most advanced systems, reportedly after Amazon flagged jailbreaks to regulators.
AWS Bedrock now requires sharing data with Anthropic and retains Mythos-class traffic for 30 days, while Microsoft has temporarily blocked employees from using Fable 5 over prompt-retention and compliance concerns.
Cybersecurity researchers say Fable’s invisible guardrails are both bypassable and overly restrictive, and the same release triggered a wave of 'we’ve reached AGI' posts even as many practitioners insist Fable 5 is still far from AGI by any rigorous definition.
In parallel, Google DeepMind published a 60‑page roadmap from AGI to ASI that explicitly defines AGI as average-human performance and ASI as surpassing large groups of experts, formalizing a target just as access to the leading contenders is being politically throttled.
open models are stealing the plot from frontier labs
DeepSeek V4 Pro, a 1.6‑trillion‑parameter open model, now beats GPT‑5.5 Pro on precision and leads coding benchmarks like SWE‑Bench and LiveCodeBench, while also being reported as dramatically cheaper per token than frontier APIs.
Qwen 3.6 27B is repeatedly cited as outperforming Gemini 2.5 Pro and Sonnet 3.7 on practical coding tasks, and the civic-backed Rio 3.5 Open 397B model from Rio de Janeiro has even outscored Qwen 3.7 on some benchmarks.
Open coding specialists like Kimi K2.7‑Code and MiMo Code show 21.8%–31.5% gains over their own predecessors on internal benches, while GLM 5.2 and MiniMax M3 bring 1M‑token context windows and multimodal capability into the open‑weights world.
Despite DeepSeek’s website share dropping to 3.4%, it already accounts for 17% of observed token volume and is driving users to downgrade expensive Claude subscriptions, hinting that the real competition is in background API calls rather than headline traffic.
token winter and the collapse of tokenmaxxing
Nvidia’s VP says compute costs now exceed employee salaries, and cases like Uber burning through its AI budget in four months are turning 'tokenmaxxing' from flex into red flag.
Enterprise users who once asked vendors for leaderboards of who could burn the most tokens are now imposing hard caps, with Meta limiting staff usage, companies shifting from tokenmaxxing to outcome selling, and reports that 90% of agent work could be replaced by simpler flows.
Subscription economics look brittle when a $200 ChatGPT plan can cost OpenAI around $14,000 if fully used and most ChatGPT‑style subscriptions are acknowledged to be heavily subsidized.
Vendors are reacting with classic wartime pricing and efficiency moves—OpenAI contemplating sharp token price cuts, Claude Fable shifting to $10 per million tokens, and techniques like speculative decoding and PoeticHQ’s 10× token savings aiming to delay an outright 'Token Winter'.
local-first stops being cosplay
Local models are no longer toys: accuracy on real‑world queries from local systems has jumped from 23.2% to 71.3%, while 'intelligence‑per‑watt' for local setups has improved by a reported 3.1×.
Google DeepMind’s Gemma 4 12B runs video, audio, and text entirely on standard laptops, Apple’s new CoreAI/MLX stack moves more inference onto Apple Silicon, and tools like MLX LM Server use continuous batching to make these local agents actually responsive.
On commodity GPUs, users report 50–80 tokens per second from local Qwen 3.6 variants via llama.cpp and similar setups, while Xiaomi and DGX Spark demos show 1,000+ tokens per second on tuned 8‑GPU servers.
Ecosystems like Ollama and LM Studio are gaining traction as secure, cloud‑independent frontends even as many would‑be users complain that the VRAM and hardware requirements for these local stacks still feel prohibitively expensive and complex.
diffusion text, video generators, and the throughput arms race
DeepMind’s DiffusionGemma abandons autoregressive text in favor of diffusion-style generation, spitting out blocks of text up to 4× faster and exceeding 1,000 tokens per second on a single H100, with only 3.8B of its parameters active at a time.
vLLM‑Omni adds speculative decoding that speeds LLMs by around 8.5× to roughly 200 tokens per second, while MiMo V2.5‑Pro and Xiaomi’s trillion‑parameter demos show 1,000‑token‑per‑second regimes on standard 8‑GPU servers.
On the generative media side, Gemini Omni Flash and Grok Imagine Video 1.5 now top text‑to‑video and image‑to‑video leaderboards, with Seedance 2.0 close behind on motion coherence, turning high‑quality video synthesis into a commodity benchmark.
But these throughput gains land in the middle of an ethics minefield, with allegations that Grok can generate about 6,000 sexual deepfakes per hour, NSFW controversies around rival systems, and a German court ruling that treats AI‑generated content as original and makes Google liable for inaccuracies.
What This Means
The center of gravity is drifting away from a single frontier chatbot toward a fractured ecosystem of politically gated Mythos‑class systems, aggressively efficient open/local models, and experimental high‑throughput architectures.
The consensus narrative of 'just wait for AGI from one lab' is quietly being replaced by a much messier story about who controls access, cost, and liability for increasingly powerful but unevenly available models.
On Watch
/OpenAI is turning ChatGPT into an agentic 'superapp' wired to Visa’s payment network and third-party services, effectively giving LLMs the ability to browse, book, and pay on users’ behalf.
/Jeff Bezos-backed Prometheus wants to build an 'Artificial General Engineer' for jet engines and spacecraft after raising $12B at a $41B valuation, explicitly betting that engineering is the next domain to get LLM‑style disruption.
/Mistral is reportedly raising about €3B at a €20B valuation around a strategy of medium-sized, fine-tuneable models and EU-compliant deployments, positioning itself as a European counterweight to U.S. labs.
Interesting
/Anthropic's retraction of a policy degrading Fable 5's performance for competitors came after significant backlash.
/Google has released Gemini-SQL2, achieving state-of-the-art results on the BIRD benchmark for translating natural language into SQL queries.
/Fine-tuning the Qwen2.5-7B model achieved 96% performance of Claude Haiku for domain-specific tasks at a low cost.
/Google's Genie 3 can turn text prompts into playable open worlds, showcasing advanced AI capabilities in game development.
/The Intelligence Frontier chart has reportedly moved backward for the first time, following the removal of certain AI models.
We processed 10,000+ comments and posts to generate this report.
AI-generated content. Verify critical information independently.
/Anthropic launched Claude Fable 5 as its first Mythos-class model, then had Fable 5 and Mythos 5 globally disabled under a U.S. export-control directive within about 72 hours.
/DeepSeek V4 Pro, a 1.6T-parameter open model, surpassed GPT‑5.5 Pro on precision and topped coding benchmarks like SWE-Bench and LiveCodeBench.
/Google DeepMind released DiffusionGemma, a diffusion-based text model that generates blocks of text up to 4× faster, exceeding 1,000 tokens per second on a single H100.
/Apple adopted a new AI architecture pairing on-device CoreAI/MLX models with cloud intelligence based on Google Gemini, including a revamped Siri.
/Kimi K2.7‑Code launched as an open-source coding model with +21.8% to +31.5% benchmark gains over K2.6 and significantly better token efficiency.
On Watch
/OpenAI is turning ChatGPT into an agentic 'superapp' wired to Visa’s payment network and third-party services, effectively giving LLMs the ability to browse, book, and pay on users’ behalf.
/Jeff Bezos-backed Prometheus wants to build an 'Artificial General Engineer' for jet engines and spacecraft after raising $12B at a $41B valuation, explicitly betting that engineering is the next domain to get LLM‑style disruption.
/Mistral is reportedly raising about €3B at a €20B valuation around a strategy of medium-sized, fine-tuneable models and EU-compliant deployments, positioning itself as a European counterweight to U.S. labs.
Interesting
/Anthropic's retraction of a policy degrading Fable 5's performance for competitors came after significant backlash.
/Google has released Gemini-SQL2, achieving state-of-the-art results on the BIRD benchmark for translating natural language into SQL queries.
/Fine-tuning the Qwen2.5-7B model achieved 96% performance of Claude Haiku for domain-specific tasks at a low cost.
/Google's Genie 3 can turn text prompts into playable open worlds, showcasing advanced AI capabilities in game development.
/The Intelligence Frontier chart has reportedly moved backward for the first time, following the removal of certain AI models.