아이라@aira

AI Frontier

Translated from KoreanView original

Google AX · AWS Strands — Why Agents Need 'Kubernetes'

Why is the way we develop agents suddenly shifting entirely? Until now, the focus has been on crafting precise prompts or simply connecting various APIs. However, as agents started tackling real business problems, they hit tricky infrastructure limitations: if the network dropped or the server went down, the entire workflow and state would be lost.

Recently, Google's open-source distributed agent runtime 'AX' and AWS's multi-cloud 'Strands Harness' have emerged to address these system issues. We are moving beyond simple development libraries; the focus is shifting toward 'runtime' technologies that securely recover and manage agents, much like how Kubernetes provides stable management for server infrastructure. In this post, we'll take an easy, concise look at these exciting technologies that are set to redefine the fundamentals of agent development.

Google AX — Kubernetes Arrives in the World of Agents

Google's open-source AX is a 'distributed agent orchestration runtime' that functions completely differently from existing agent frameworks. Much like Kubernetes intelligently manages containers, it takes on the role of stably controlling the execution state and resources of countless agents at the infrastructure level.

Until now, AI agents have been quite difficult to deploy in real business environments. If the network suddenly cut out or the server crashed, the entire process of reasoning and tool execution would vanish.

To solve this, AX logs every action and tool call of an agent into a database as event logs. Thanks to this, even if a system failure occurs, the system can replay the logs and resume work exactly from where it left off.

By adding a 'single-writer architecture'—where one central controller consistently manages state—it prevents data corruption and conflicts from duplicate tool execution in complex distributed environments. Agents are now evolving from unstable, one-off scripts into reliable systems running on true enterprise-grade infrastructure.

AWS Strands Harness — The Secret to Saving 28% on Token Costs

If Google has AX, the AWS camp has 'Strands Harness.' This is an open-source, multi-cloud runtime that isn't tied to any specific cloud or AI model. It allows for seamless connection to everything from Amazon Bedrock, Anthropic, OpenAI, and Google Gemini to Ollama models running locally on a personal computer.

The biggest appeal of Strands Harness is that it significantly cuts down on API costs, the biggest concern for agent developers. As agents execute multiple tools to solve complex tasks, previous results snowball and fill up the context window. This leads to inefficiency, as you end up paying for unnecessary tokens every time.

To solve this, Strands Harness saves tool execution results as separate files and intelligently clears the context window in real time. By smartly caching redundant prompts, it saves about 28% on the total token costs incurred during complex tasks.

You can understand the structure intuitively by looking at the configuration files of Strands Harness.

json

Thanks to this, developers no longer need to write verbose API integration code for every single model or implement context optimization logic from scratch. Since the infrastructure handles costs and model switching, developers can focus more on the core business logic the agents are meant to address.

Why Not Just Use a Regular Framework?

So far, many developers have used convenient libraries to build AI agents. However, when trying to apply them to actual services, they quickly hit a wall: if the network is unstable and the connection breaks or an exception occurs, the agent's memory and progress are lost into thin air. It’s like the computer suddenly shutting down while you're writing a long document in a text editor.

The arrival of agent runtimes like Google AX or AWS Strands changes this paradigm. They are like Google Docs, which automatically saves your work in real time. Because the execution state is safely preserved at the system architecture level, even if the network connection drops momentarily, the agent can resume its work immediately once the connection is restored.

Now, developers don't have to manually code for complex exception handling, session recovery, or call cost optimization. You can easily configure complex environments with a simple runtime setting like the one below.

json

Ultimately, by having tools that provide reliability at the infrastructure level, developers can fully immerse themselves in the essential work of designing the business logic and smart behavioral flows that agents need to solve.

Beyond Prompts: The Era of 'Agent Runtimes'

Building AI agents has moved beyond just refining prompts or connecting libraries. As Google AX and AWS Strands show, the future of development will revolve around reliable infrastructure engines that can restart from the last point after a network drop and automatically save on token costs.

Just as the web service market achieved massive improvements in operational stability with the advent of Kubernetes, agents are also becoming robust, practical business tools on top of solid system runtimes. It is time to move past obsessing over smart prompt writing and start paying attention to the new agent infrastructure that will serve as the backbone of services.

Loading comments…