TL;DR
Grok just took the coding crown while simultaneously leaking people’s directories, orchestration layers started mattering more than whichever frontier model you pick, and serious users are quietly retreating to local, small models on cheap GPUs and even phones.
The real frontier now lives in the stack—routers, runtimes, and security posture—while AGI talk floats far above today’s very physical bottlenecks in energy, memory, and data.
Key Events
Report
The headline story isn’t that models got smarter; it’s that where and how you run them suddenly matters more than which one you pick. Coding SOTA just flipped to Grok 4.5 while orchestration layers, local runtimes, and tiny models quietly stole most of the real gains.
On coding-style tasks, Grok 4.5 now beats Claude Fable 5 and GPT-5.6 Sol on SWE-Atlas-QnA, even as Sol still holds an 80.0 on the Artificial Analysis Coding Agent Index.
Debate-style leaderboards put Muse Spark 1.1 just behind Claude’s best and ahead of Opus on AA-Briefcase and the Debate Benchmark. Multimodal text-to-video is led by Gemini 3.5 / Omni Flash, while open models like Wan-Dancer and LingBot-World-Infinity push long-horizon video far beyond meme generators.
Meanwhile, a new 8B model and compressed Qwen3.5-4B / gemma-4-12b variants are reported beating GPT‑4–scale systems on some complex or reasoning-trace tasks.
The net effect is that “best model” discourse has collapsed into a patchwork of task-specific winners rather than a single frontier idol.
Research-grade orchestrators juggling around six foundation models report 33–61% cost cuts and 44% median latency reductions while keeping quality flat, and production stacks are already hot-swapping between Fable 5, GPT‑5.6 and others.
GitHub Copilot’s CLI improved subagent delegation to reduce tool failures, while Codex’s “Ultra” subagents quietly balloon token usage through nested spawns.
At the same time, MCP and tool-rich agents are hitting hard limits: tool-selection accuracy collapses past ~20 tools, and misconfigured agents can loop on failing APIs until they exhaust budgets.
Security is catching up only reactively—MCP servers let agents call any exposed tool, Grok Build’s CLI uploaded entire Git repos and env files, and users are bolting on default-deny firewalls to things like OpenClaw and Hermes after the fact.
PrismML has a compressed Qwen 3.6‑27B running on an iPhone 17 Pro, while GLM 5.2 does 2–2.8 tokens/s on a MacBook Pro M5 and still claims wins over GPT‑4 on coding.
Gemma 4 runs entirely inside Godot via GDScript and Vulkan, powering autonomous local NPCs, and its 5:1 local:global attention ratio shrinks KV cache footprints.
On the “boring” side, DeepSeek v4 Flash is now the go-to local Excel model at 30–40 tok/s, and Nobie wraps an Excel-compatible runtime around humans and agents instead of trying to rip spreadsheets out of workflows.
Used enterprise GPUs like P100s and V100s cost under $200, Colibri streaming runs models at 10GB RAM, and community favorites like llama.cpp make 20–50B models usable on refurbished desktops.
An automated governance monitor watched 6,200 hours of footage across 47 AI summits and found zero enforceable commitments, while real incidents kept landing on Twitter instead of in regulators’ inboxes.
Grok CLI and Build uploaded entire home directories, Git repos, and environment files to Google Cloud until a hidden server flag partially stopped it, and Zero Data Retention mode is gated behind enterprise accounts.
OpenAI’s personal plans have data sharing enabled by default, MCP servers routinely expose broad tool surfaces with shaky access control, and third-party storage systems like ShareFile are shipping urgent security alerts.
Power users are responding by fleeing to local models, default-deny firewalls, and self-hosted homelabs built on refurbished desktops and small LLM runtimes.
Richard Sutton’s new Oak Lab is explicitly targeting a trillion‑parameter agent that learns and plans in real time on only 20 watts, and forecasting models are calling for fully automated coding on AGI projects by June 2028.
ASI discourse layers on top, promising objective morality, immortality tech, and a fully automated economy that ends most human suffering.
Yet the present tense looks much less godlike: Meta’s AI data center spend jumped from ~$10B to ~$50B in under two years, electricity demand from data centers is set to climb 26% this year, and Ireland’s data centers will soon match all homes’ power use.
On the model side, GPT‑5.6 Sol is having its “juice” budgets cut for efficiency, clusters sit idle because of storage constraints, and practitioners are grinding on dataset quality, LoRA training, and evaluation tricks like Instance, ReContext, and J-space entropy.
The disconnect between trillion‑parameter, 20‑watt dreams and very physical bottlenecks in energy, memory, and data is where the real AGI story now lives.
What This Means
Across all of this, the center of gravity is sliding from “pick the smartest model” to “compose a trustworthy, efficient stack” where benchmarks, orchestration, local hardware, and governance gaps all matter as much as raw IQ. The wild part is that the system-level frontier now lives in glue code, runtimes, and security posture rather than in single model weights.
On Watch
Interesting
We processed 10,000+ comments and posts to generate this report.
AI-generated content. Verify critical information independently.
Sources
Key Events
On Watch
Interesting