Back to Blog
Research

The AI Pricing Floor Just Collapsed. Here's What $0.20/Million Tokens Actually Changes for Your Stack

The Workflow Finder
2026-08-02
3 min read
The AI Pricing Floor Just Collapsed. Here's What $0.20/Million Tokens Actually Changes for Your Stack

OpenAI cut GPT-5.6 Luna's price 80% in three weeks. The real story isn't the discount, it's which agentic workflows just became affordable.

On July 30, OpenAI cut the price of GPT-5.6 Luna by 80%, from $1/$6 to $0.20/$1.20 per million input/output tokens. Terra dropped 20%, from $2.50/$15 to $2/$12. Sol held steady. Price cuts happen constantly in this industry and most of them don't matter enough to write about. This one does, because of what OpenAI said caused it and what it actually unlocks downstream.

The model cut its own hosting cost

OpenAI's own explanation is the interesting part: GPT-5.6 Sol reportedly rewrote production GPU kernels to cut serving cost by roughly 20% and improved speculative decoding by more than 15%. That's a model making itself cheaper to run, then passing the savings to the cheaper model in the same family. Whether or not you take the framing at face value, the resulting price is real and independently confirmed by CNBC and OpenAI's own developer forum. Luna at roughly $1.40 per million tokens combined now undercuts Google's Gemini 3.5 Flash-Lite ($2.80/M combined) while reportedly scoring higher on Artificial Analysis's intelligence index, according to VentureBeat's reporting. An OpenAI-branded model being the cheap option instead of the premium one is new.

Why this is a competitive move, not generosity

CNBC's framing on this is blunt: the cuts are a direct reaction to Chinese open-weight pressure, specifically Kimi K3's July 16 release and the broader pattern of frontier-adjacent benchmarks showing up at sub-$5/M combined pricing from multiple Chinese labs. Google had been citing Gemini 3.6 Flash's $9/M combined price as cheaper-per-task than the Chinese alternatives. OpenAI's cut directly undercuts that argument. This is a price war with three active participants, not a single vendor being generous.

What actually changes at these prices

The math that matters is agentic workload math, not chatbot math. A chatbot conversation might use a few thousand tokens. An autonomous multi-hour agent run can burn millions. At Luna's old $7/M combined rate, a 5-million-token agent session cost $35. At the new $1.40/M combined rate, the same session costs $7. That's not a rounding difference, it's the difference between "we can't afford to run this at scale" and "run it." Document summarization pipelines, lightweight classification layers, real-time assistant backends, and multi-step agent orchestration all become economical at volumes that didn't make sense three weeks ago. OpenAI also quietly upgraded Codex CLI's Auto-review step from GPT-5.4 to Luna, making that safety check roughly ten times cheaper to run without anyone having to change a line of code.

The catch: speed is now the premium tier

Sol didn't get cheaper, it got a new Fast mode instead, 2.5x the speed at 2x the price ($10/$60 per million tokens), replacing the old Priority Processing option. Existing requests tagged priority route to Fast mode automatically. Read that as the actual shape of the new pricing structure: the cheap tier is racing to the floor on cost, and the premium tier is racing up-market on latency instead of intelligence. If your workload needs Sol-level reasoning fast, that combination is now explicitly what you're paying for, separately from the reasoning itself.

What to actually do with this

If you built anything on GPT-5.6 Terra or Luna before July 30 and didn't revisit the numbers, go check your bill, you're probably already paying less without having changed anything. If you ruled out an agentic workflow six weeks ago because the token math didn't close, run the numbers again at the new Luna rate before assuming it's still dead. And if you're evaluating vendors right now, price is moving fast enough in this category that whatever number is on a pricing page today is a snapshot, not a commitment, check it again before you finalize a build that depends on it.

Mentioned in This Post

Share this article

Related articles

Signal, no noise.

A weekly breakdown of the AI tools and workflows actually worth your time.