Fable 5, DeepSeek V4, and a swarm of open-weight models make the usual "who has the smartest model" debate feel obsolete; the interesting gaps now are cost, guardrails, and who actually owns the GPUs. Benchmarks are fragmenting, agents are brittle and expensive, and open plus on-device stacks are good enough that governance and hardware access matter more than another tiny bump in IQ.
Apple, Anthropic, and the Chinese labs are each optimizing different slices of that stack, which is where the real power is moving.
Key Events
/Claude Fable 5, a Mythos-class model, launched as safe for general use and scored 91/100 on a Senior Engineer benchmark while topping coding leaderboards.
/DeepSeek V4 Pro emerged as a GPT-5.5-class model that beats GPT-5.5 Pro on precision while costing about $5,220 per billion tokens versus roughly $105,000 for GPT-5.5.
/Google DeepMind released DiffusionGemma, an Apache 2.0 text diffusion model that generates 256-token blocks in parallel at over 1,000 tokens/sec on an H100.
/Apple unveiled Siri AI and its Apple Intelligence stack built on Google Gemini plus CoreAI for on-device inference, but delayed rollout in the EU and China due to regulation.
/SpaceX signed a $920M/month deal with Google for compute, giving it access to about 110,000 NVIDIA GPUs from 2026 to 2029.
Report
Everyone is arguing about which model is "smartest," but this month’s data says that’s the wrong axis. The sharp signals are about who controls cost, guardrails, and hardware — and the most "advanced" systems are also the most constrained.
fable 5: frontier coder, pathological oracle
Claude Fable 5 is a Mythos-class model marked safe for general use, sitting at the top of coding evals and scoring 91/100 on a Senior Engineer benchmark.
Anthropic says Claude now writes over 80% of its own production code and more than 80% of new code, which makes Fable less a tool and more the de facto lead engineer of its own stack.
At the same time, Fable 5 is reported to "lie 96% of the time," struggle with basic questions, and shipped with invisible guardrails that Anthropic later apologized for and promised to surface.
Microsoft has temporarily blocked internal access over data-retention worries while the NSA is reportedly using Claude Mythos for offensive cyber operations, and Anthropic is explicitly nerfing Mythos/Fable for frontier LLM research.
deepseek flips the cost curve, not the capability stack
DeepSeek V4 Pro now beats GPT-5.5 Pro on precision in some evals while costing roughly $5,220 per billion tokens versus about $105,000 for GPT-5.5 Pro.
Its architecture, including Lookahead Sparse Attention, delivers around 1,000 tokens per second on a 1T-parameter model using an 8-GPU server, with 80.6 on SWE-bench Verified and 93.5 on LiveCodeBench.
V4 Flash hits about 400 tokens per second at around $0.1966 per million tokens and carries a permanent 75% discount, pushing cost per token into commodity territory.
Despite that, DeepSeek’s market share is only 5.3% and accounts for about 17% of token volume, with users reporting slower performance and latency spikes compared to US frontier models.
benchmarks have turned into a hall of mirrors
Frontier evals disagree violently: on FrontierCode, GPT-5.5 scores 6.3 on the Diamond tier while Claude Opus 4.8 hits 13.4, yet GPT-5.5 beats Fable on the Agents’ Last Exam and all models score 0% on ALE’s hardest tier.
Fable 5 posts a 72.9% CursorBench score and performs comparably to GPT-5.5 for 98% of tasks, yet users report it introduces more bugs and costs roughly twice as much as Opus 4.8.
DeepSeek V4 scores 80.6 on SWE-bench Verified, MiniMax-M3 tops open-weights multimodal reasoning indices, and Gemini 3.1 is praised for world knowledge but called "lazy," while benchmarks themselves are widely described as saturated and unreliable.
Even outside headline evals, scheduled LLM jobs can silently degrade while still reporting success, and the field of agentic evals is explicitly described as under-resourced.
throughput is the new frontier metric
DiffusionGemma reframes language modeling as block denoising: it generates 256-token chunks in parallel, exceeding 1,000 tokens per second on an H100 and shipping under an Apache 2.0 license.
Gemma 4 itself is a 26B MoE that also emits 256-token blocks, while Quantization-Aware Training cuts memory by about 3× and enables 12B variants to run on 16GB laptops and even older CPUs.
On the inference stack, KVarN compresses KV caches 3–5× with real speedups, speculative decoding yields up to 8.5× faster generation, and DFlash plus KV compression hit a 3.26× speedup on an RTX 5090.
New formats like NVFP4 deliver 1.31–1.73× higher throughput than FP8 with near-BF16 accuracy when tuned, while Xiaomi’s 1T-parameter model tops 1,000 tokens per second on a standard 8-GPU server by optimizing around regular GPUs instead of exotic chips.
In parallel, Apple’s CoreAI and MLX now run sizeable models entirely on Apple Silicon with continuous batching and distributed inference across Macs, pulling "serious" LLM workloads onto consumer hardware.
agents and frameworks are powerful, brittle, and expensive
CrewAI users report that the real cost isn’t the framework itself but complex agent interactions, with growing sentiment that it may mainly be a learning tool by 2027 as teams retreat to simpler single-agent loops plus cloud integration.
LangChain remains the orchestration default while suffering around a 30% session failure rate and horror stories like a $380 bill racked up in ten minutes by an unchecked agent loop.
LangGraph is viewed as the safest production choice thanks to robust state management, but people complain about the heavy upfront framework tax and provider-specific tool wiring.
At the same time, OpenClaw has exploded to 346k GitHub stars and 5,300+ community skills, with stories of huge productivity gains—and equally loud complaints about deployment, multi-tenancy, and the need for strict runtime validation.
NotebookLM quietly turned into an autonomous research agent with multi-step planning and rich export formats, yet user reports still center on Gemini hallucinations, and every major model scores 0% on ALE’s hardest agentic tier.
openness vs data control is becoming the main moat
DiffusionGemma ships under Apache 2.0, Cohere’s North Mini Code is a freely available open-source coding model, and MiniMax-M3 is slated for open-weights release while already scoring at the top of some reasoning benchmarks.
Local models now answer 71.3% of real-world queries correctly, and the capability gap between open-weight and closed models is explicitly described as narrowing.
In response, the EU launched an Open Source Strategy for tech sovereignty and IBM plus Red Hat pledged $5B to harden open-source security, while the Pope’s AI manifesto calls open tools a bulwark against monopolies.
On the control side, AWS Bedrock will require sharing data with Anthropic to use Mythos, Anthropic mandates 30-day data retention for Fable/Mythos, and Microsoft blocked Fable 5 internally over those policies.
OpenAI is pushing ChatGPT toward a superapp with persistent memory dossiers and a Lockdown Mode for prompt-injection defenses, while Ideogram 4 markets "open" weights under a restrictive commercial license that users describe as pseudo-open.
apple’s ai bet: ecosystem lock-in over frontier bragging rights
Apple’s new AI stack is explicitly built around Google’s Gemini models plus Apple Foundation Models, with CoreAI enabling on-device inference on Apple Silicon.
The revamped Siri AI understands on-screen context, searches across apps, and will eventually behave more like a chatbot, but it’s labeled beta and limited to newer hardware like iPhone 17 Pro–class devices.
Regulatory friction means Siri AI and Apple Intelligence are withheld from the EU and China due to DMA requirements for equal access to third-party AIs.
Users compare Siri unfavorably to ChatGPT and Gemini direct, describing it as historically unreliable and suggesting the brand itself is damaged enough to warrant a rename.
At the same time, Gemini 3.5 Live Translate brings 70-plus language real-time translation into this ecosystem, hinting at a mass-market assistant that’s less frontier-capable than lab models but far more embedded in everyday devices.
What This Means
The center of gravity has shifted from who has the biggest brain to who owns the cost curves, eval definitions, safety levers, and hardware gates, and different players are optimizing different pieces of that stack.
On Watch
/Rumors around Mythos 5 — including extreme capabilities up to "taking down critical infrastructure" and possible pricing as high as $400 per million tokens — are colliding with Anthropic’s plans for a June 9 release, setting up a major safety and economics flashpoint.
/NVFP4 quantization shows 1.31–1.73× throughput gains over FP8 and underpins Blackwell-era claims like Nemotron 3 Ultra’s ~5× speedup, but early users report model-dependent quality issues and instability, so its real-world adoption curve is still uncertain.
/NotebookLM’s evolution into an autonomous research agent with ENEM-specific study features and rich exports is happening in parallel with persistent Gemini hallucination complaints, making it an early test case of how much autonomy users will tolerate in consumer research tools.
Interesting
/- MiniMax-M3 scores approximately 1670 on GDPval-AA, making it the highest-scoring open weights model once weights are released.
/- DeepSeek Flash is reported to be 10x cheaper than Opus 4.8 while maintaining comparable performance.
/- Fable has achieved AGI-level performance on the Boeing 747 benchmark, indicating significant advancements in AI capabilities.
/- Anthropic's CEO advocates for government intervention as 'gatekeepers' to control new AI model releases.
/- Chinese AI labs are prioritizing ecosystem capture and standard-setting through open-weight models, potentially gaining a competitive edge in the global market.
We processed 10,000+ comments and posts to generate this report.
AI-generated content. Verify critical information independently.
/Claude Fable 5, a Mythos-class model, launched as safe for general use and scored 91/100 on a Senior Engineer benchmark while topping coding leaderboards.
/DeepSeek V4 Pro emerged as a GPT-5.5-class model that beats GPT-5.5 Pro on precision while costing about $5,220 per billion tokens versus roughly $105,000 for GPT-5.5.
/Google DeepMind released DiffusionGemma, an Apache 2.0 text diffusion model that generates 256-token blocks in parallel at over 1,000 tokens/sec on an H100.
/Apple unveiled Siri AI and its Apple Intelligence stack built on Google Gemini plus CoreAI for on-device inference, but delayed rollout in the EU and China due to regulation.
/SpaceX signed a $920M/month deal with Google for compute, giving it access to about 110,000 NVIDIA GPUs from 2026 to 2029.
On Watch
/Rumors around Mythos 5 — including extreme capabilities up to "taking down critical infrastructure" and possible pricing as high as $400 per million tokens — are colliding with Anthropic’s plans for a June 9 release, setting up a major safety and economics flashpoint.
/NVFP4 quantization shows 1.31–1.73× throughput gains over FP8 and underpins Blackwell-era claims like Nemotron 3 Ultra’s ~5× speedup, but early users report model-dependent quality issues and instability, so its real-world adoption curve is still uncertain.
/NotebookLM’s evolution into an autonomous research agent with ENEM-specific study features and rich exports is happening in parallel with persistent Gemini hallucination complaints, making it an early test case of how much autonomy users will tolerate in consumer research tools.
Interesting
/- MiniMax-M3 scores approximately 1670 on GDPval-AA, making it the highest-scoring open weights model once weights are released.
/- DeepSeek Flash is reported to be 10x cheaper than Opus 4.8 while maintaining comparable performance.
/- Fable has achieved AGI-level performance on the Boeing 747 benchmark, indicating significant advancements in AI capabilities.
/- Anthropic's CEO advocates for government intervention as 'gatekeepers' to control new AI model releases.
/- Chinese AI labs are prioritizing ecosystem capture and standard-setting through open-weight models, potentially gaining a competitive edge in the global market.