TL;DR
Money is rotating down the AI stack: HBM and custom chips are printing cash while generic model APIs head toward a price war. At the same time, top-end models like Mythos and GPT‑5.6 are being treated less like software and more like controlled weapons, even as China and open source drive inference costs toward zero.
The hard part now is separating the few real infra moats from the many AI bubble stories riding their coattails.
Key Events
Report
HBM vendors and custom silicon makers are where the real AI money is accruing right now, not the headline GPU brands. Governments are also starting to treat top‑end models as regulated weapons rather than generic software.
Micron’s stock is up nearly 800% over the past year on an AI‑driven memory crunch. The same demand has effectively quadrupled Micron’s revenue and let it secure historically high DRAM prices on long‑dated contracts.
SK Hynix has just overtaken Samsung as South Korea’s most valuable company on the back of its high‑bandwidth memory leadership and is even reallocating some HBM capacity back to general‑purpose DRAM.
Chinese challenger CXMT is projected to reach about $55B of revenue and both SK Hynix and Samsung are planning aggressive capacity expansions, setting up a classic cycle between near‑term scarcity and eventual glut risk.
OpenAI’s first custom chip, Jalapeño, co‑developed with Broadcom and built at TSMC, is marketed as state‑of‑the‑art in performance per watt for LLM inference workloads.
Qualcomm is paying $4B for AI software firm Modular and is in talks to acquire Tenstorrent, explicitly to bulk up its AI software and accelerator stack.
Google is spending billions to turn its TPUs into a credible Nvidia challenger, with Citadel reporting key workloads running 30% cheaper and up to four times faster on TPUs than on GPUs.
Nvidia’s market cap is about $5.05T, making it the most valuable public company. It also signed a $6.3B GB300 GPU supply deal with Reflection AI and SpaceX running through 2029, showing big buyers will still pre‑commit to long‑dated capacity even as alternatives appear.
Nvidia‑class AI data centers are estimated at about $47B per gigawatt of capacity with annual electricity bills around $1.3B per GW, before headcount.
Oracle has already laid off roughly 21,000 people in 12 months, explicitly tying part of the cuts to AI adoption and expensive infrastructure bets, and is facing questions about whether AI data‑center loans are sustainable.
Meta cut 8,000 jobs right after a record $56.3B quarter, while six executives received options worth up to $921M each and internal leaders describe morale as “the worst it’s ever been.” Amazon, Meta, Walmart and Uber are now putting hard caps on internal AI usage because of budget strain, a sharp reversal from last year’s “AI everywhere” stance.
At the same time, generative AI has produced about $110B in sales over 12 months amid increasingly loud bubble talk, including dot‑com comparisons, forecasts of a $1T tech sell‑off, and reports that OpenAI has already missed revenue targets and needs substantial new capital.
Anthropic’s Mythos model reportedly breached almost all NSA classified systems within hours during a red‑team test, after which the US government banned it as “too dangerous” while still allowing around 200 companies to keep access.
The same government has ordered a staggered, customer‑by‑customer rollout of GPT‑5.6 and is preparing a formal licensing regime for access, putting frontier models into a controlled‑technology bucket.
Anthropic is rolling out identity verification for sensitive capabilities from July 2026 and has already cut off access to its models for many foreign nationals under export‑control directives.
The EU AI Act will require watermarking of AI‑generated text from August 2 with significant fines for non‑compliance, and Big Tech has already eaten a $3.5B penalty for using personal data to train AI.
Norway has imposed an effective ban on AI in elementary schools and Amnesty International has labelled leading generative systems “unlawful by design,” while only 16% of Americans tell pollsters they expect AI to benefit society.
Five major Chinese AI labs, including Alibaba, ByteDance and Tencent, have slashed token prices by between 50% and 99% as capability gaps between their models narrow.
Chinese resellers are offering Claude access at 70–90% below official pricing through fraudulent token reselling into blocked markets. Anthropic separately claims Alibaba used nearly 25,000 fake accounts to run 29M Claude exchanges to siphon capabilities for its Qwen lab.
In parallel, open‑weight models are closing the performance gap: GLM‑5.2 from Beijing’s Z.ai ranks near the top of new agentic benchmarks, can run locally via llama.cpp after aggressive size reduction, and is reported to beat GPT‑5.5 on several tasks at a fraction of the price.
The EU is funding a 400B‑parameter open‑source frontier model, EUROPA, on European supercomputers, while DDR5 price drops are explicitly boosting local LLM building and Krea 2 and other open‑weight systems show up as credible text‑to‑image alternatives.
What This Means
Power in AI is sliding down the stack to whoever controls scarce components, custom silicon and regulatory relationships, while generic model access is being commoditized and politically constrained. The live uncertainty is which parts of today’s AI infra and model layer are durable profit pools versus transient bubble stories.
On Watch
Interesting
We processed 10,000+ comments and posts to generate this report.
AI-generated content. Verify critical information independently.
Sources
Key Events
On Watch
Interesting