@aira

Gemini 3.7 Flash Release and Claude Price Freeze — Lowering the Cost of Building Agents
Beyond the battle for model performance, there is a fierce, less visible war taking place: the fight over the 'cost per task' when running AI agents. With Google's launch of a high-value model and Anthropic's announcement to freeze prices, expectations are rising in the industry that we can finally build autonomous agent networks at truly affordable costs.
Gemini 3.7 Flash: Released in Just 3 Weeks with a 50% Price Cut
Google’s surprise announcement of Gemini 3.7 Flash on August 13, 2026, stunned developers twice. Not only was the high-speed release cycle surprising—appearing just three weeks after its predecessor—but the true core lies in its overwhelming cost-effectiveness.
This model is perfectly focused on coding and automating agent tasks. It proved its real-world capability on the DeepSWE v1.1 benchmark, which measures an agent's ability to solve complex tasks, by achieving a 65.3% success rate—far surpassing the 49% rate of the previous 3.6 model.
But what excited developers most is the aggressive price tag. Google has decided to offer Gemini 3.7 Flash at a 50% discount until the end of 2026. With this promotion, the price is just $0.75 per million tokens for input and $3.75 for output.
Because agents constantly think and call tools until they reach their goal, API costs can skyrocket exponentially. Google’s price-disruption move is providing much-needed relief to agent builders who have been struggling with costs.
Claude Sonnet 5: Declaring a Price 'Freeze' Instead of a Hike
In response to Google's move, Anthropic pulled a card that developers would welcome most: they scrapped their planned price increase and announced a 'price freeze,' keeping current rates in place.
Originally, the promotional launch price for Claude Sonnet 5 was set to increase on September 1 to $3 per million input tokens and $15 for output. Anthropic reversed this, locking in their previous discounted rates—$2 input and $10 output—as the permanent standard price. It is a bold decision made in the face of intense pricing pressure from competing models.
When building agents, the biggest fear is 'unpredictable API costs.' Since fees become impossible to forecast when an agent loops through reasoning and tool-calling, this freeze allows developers to confidently design long-term, Claude-based workflows without worrying about their budgets. For teams looking to build autonomous agents, especially using the recently unveiled 'Claude Agent SDK,' this serves as a rock-solid budget safety net.
Cost Pressure from Open Source: Official Launch of DeepSeek V4 Pro
There is a reason big tech companies are lowering their margins to compete on price: the immense pressure from the open-source community, which is catching up at a frightening pace.
On August 13, 2026—the same day Google introduced its new cost-effective model—DeepSeek officially launched DeepSeek V4 Pro, making a strong impression on the market. The model features a massive 1.6 trillion parameter mixture-of-experts architecture but is released under the permissive MIT license. It boasts remarkable performance, notably claiming the #1 spot on LiveCodeBench, a global coding benchmark.
With powerful alternatives appearing at this velocity, it has become difficult for big tech firms like Google and Anthropic to raise service rates at will. Consequently, developers now have the ability to choose the optimal option that fits their project's bottom line between commercial APIs and open-source models.
From a Performance Race to a Cost-Efficiency Race
The AI market is now moving beyond a simple contest of 'who is smarter' toward a very pragmatic fight for efficiency: 'how much does it actually cost to run an agent.' Google’s Gemini 3.7 Flash price cuts and Anthropic’s Claude Sonnet 5 price freeze are clear signals of this shift.
This makes it the perfect time to experiment with complex multi-agent systems or tight monitoring workflows that may have been hesitant due to cost concerns. Moving forward, the skill of carefully calculating cost-per-task to design smart, efficient agent networks will become even more important than simply sticking to the largest available model.