The cost of thinking, at least the artificial kind, just got a lot cheaper. Average prices for AI inference from leading US labs dropped nearly 25% between mid-July and mid-August, according to Silicon Data analysis reported by the Financial Times.
Who cut what, and by how much
OpenAI moved first on July 30, taking a machete to pricing on two of its three GPT-5.6 model tiers. GPT-5.6 Luna, the company’s mid-range offering, saw an 80% price reduction. GPT-5.6 Terra got a 20% haircut. The flagship GPT-5.6 Sol held steady, though OpenAI sweetened the deal by accelerating its performance options at the same price point.
Anthropic followed a similar playbook, launching Claude Opus 5 at half the price of its predecessor, Fable 5.
The DeepSeek effect
The pressure is coming primarily from Chinese providers, with DeepSeek and Moonshot leading the charge. These firms offer models that either match or closely approach US performance benchmarks, but at dramatically lower price points.
To put the broader trajectory in perspective: GPT-4-class performance cost over $20 per million tokens in late 2022. By mid-2026, that same tier of capability costs less than $1 per million tokens across various providers. That’s a decline of more than 95% in roughly three and a half years.
Good for users, complicated for investors
For enterprises that have been cautiously experimenting with AI, cheaper inference is unambiguously good news. Lower costs remove one of the biggest barriers to broader adoption, letting companies run more experiments, deploy more agents, and integrate AI into workflows that previously didn’t pencil out at higher price points.
US AI labs have attracted enormous valuations built partly on the expectation that inference would remain a high-margin business. When your revenue-per-query drops by 25% in a month, the financial models that justified those valuations need revision. OpenAI’s decision to hold GPT-5.6 Sol’s price while cutting Luna and Terra suggests it’s betting on a tiered strategy, where the best model keeps its margins while cheaper tiers fight for volume.
US labs have spent billions building out inference infrastructure. Those investments were justified by revenue projections that assumed certain price levels. When prices fall faster than usage grows, the return on that infrastructure spending gets stretched, potentially reshaping how investors evaluate the entire AI supply chain from chips to cloud providers to the labs themselves.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
7








English (US) ·