아이라@aira
AI Frontier

Claude Sonnet 5.5 Released — Terminal Coding Success Rate Surpasses 70%
Anthropic has launched the new AI model 'Claude Sonnet 5.5' without prior notice. Following the release of their flagship Claude Opus 3.5 just a week ago, they have dropped another powerful card in quick succession. The new Sonnet 5.5 is not only over 30% faster but has significantly boosted its 'autonomous agent' capabilities, allowing it to execute terminal commands and modify code on its own. Let's take a quick look at why the developer community is so hyped about this new model.
From 10% to 70% Terminal Success: What Has Changed?
The most eye-catching change comes from the Terminal-Bench 4.0 evaluation. This benchmark is a rigorous test that moves beyond simple code generation to verify whether an AI can complete tasks by entering commands directly into a real computer terminal environment.
In this test, the previous generation, Claude Sonnet 3.5, had a terminal task success rate of only 10.3%. It struggled to manipulate complex system commands, making it quite limited as an autonomous agent.
However, the newly released Sonnet 5.5 achieved a record-breaking score of 70.6%. This is clear evidence that the AI has moved past basic code suggestions and can now reliably act as an autonomous coding agent, installing necessary packages and debugging errors in virtual environments.
Adaptive Thinking — API for Controlling Reasoning Depth
One of the most interesting updates in Claude Sonnet 5.5 is 'Adaptive Thinking,' which allows the model to control its own internal reasoning process before answering. Developers can now set the effort level in API calls to determine how deeply the model should think to solve a problem. For example, you can lower the effort level for light text processing to get quick answers, or increase it for complex system architecture or security audits to ensure precision.
It is important to note that the disable setting used to completely turn off thinking in the previous Claude Sonnet 3.5 is gone. Starting with Sonnet 5.5, a new 'thinking-between-tools' option has been introduced to prevent unnecessary pre-reasoning while keeping the agent loop intact. Enabling this option skips the long initial reasoning phase to save time and costs, while still providing visibility into the progress as tools are executed.
Here is an example configuration based on the Claude Sonnet 5.5 API.
With this ability to flexibly control the depth of thought via the API, development teams can prevent unnecessary token waste and allocate resources efficiently based on the specific use case. It effectively provides a powerful control mechanism to help autonomous agents act intelligently depending on the situation.
Same Price, Faster Speed: How Does It Stack Up on Value?
Anthropic has set the base API price for Claude Sonnet 5.5 identical to the previous Sonnet 3.5. At $2 per million tokens for input and $10 for output, the fact that they haven't raised prices despite such significant performance gains is great news for developers.
Beyond the price freeze, the actual cost of completing tasks has effectively dropped by up to 30%. Since Sonnet 5.5 runs over 30% faster than the previous version and minimizes token waste by reducing unnecessary tool calls, you can get the same work done faster and with fewer tokens.
However, keep in mind that if you utilize the new Adaptive Thinking feature and set the reasoning depth to maximum, it will consume more tokens during the inference process. Therefore, matching the effort level to the complexity of the task will be the key to efficient development cost management.
Handling Terminals with Ease: A True Collaborative Development Partner
The arrival of Claude Sonnet 5.5 means more than just having another smart autocomplete tool. AI is evolving into a reliable fellow engineer that can enter commands directly into the terminal, fix errors in real-time, and produce results.
The synergy is even greater when combined with the Model Context Protocol (MCP), an open-source standard that connects various development tools and data. Tedious tasks that developers used to do manually, like executing terminal commands, are now part of a workflow the AI handles itself.
Development in the future will likely involve AI agents leading the tasks while humans focus on final reviews and approvals. With Sonnet 5.5 having cleared the barrier of terminal interaction, we encourage you to experience a more convenient and creative development process firsthand.