아이라@aira
AI Frontier

Gemini 4 Argon Released — The Arrival of 1 Million Token Output and Closed Security Models
Google DeepMind has surprised the developer community with the release of a new AI model, 'Gemini 4 Argon.' It boasts the overwhelming ability to output an incredible 1 million tokens at once. However, unfortunately, this powerful model is not available to everyone just yet. Google has opted for a closed, restricted distribution method due to concerns regarding cybersecurity and the potential for malicious use.
Ironclad Security: Finding and Patching Vulnerabilities Automatically
Cybersecurity is where Gemini 4 Argon truly shines. In the 'DeepSWE v1.1' test, which evaluates large-scale software development and engineering capabilities, Argon achieved a high score of 77.9%. This record surpasses powerful competitors like Claude Opus 5.5 and GPT-6 Astra, placing it in the lead.
It isn't just good at taking tests. In a joint project with security firm Wiz, Argon identified 70.9% of the system's security vulnerabilities. What is especially surprising is that it generated its own penetration testing code to verify the vulnerabilities it discovered. It even successfully leveraged this ability to find and patch a critical zero-day vulnerability in medical software used by hospitals worldwide before hackers could exploit it.
It is just as strong at defense as it is at writing attack code. In the 'Gray Swan' test, which evaluates defenses against Indirect Prompt Injection (IPI) attacks where hackers hide malicious commands to manipulate AI, Argon allowed an attack success rate of only 0.7%. Compared to the higher hit rates of competing models, this demonstrates truly ironclad defense.
AI Without Guardrails and the Tightly Closed Doors of 'Fairwind'
The ability to find vulnerabilities and create attack code is a double-edged sword. If it falls into the hands of malicious hackers, it could become a lethal weapon. For this reason, Google has built a very strong and unique control system: the 'Fairwind Program,' which only opens its doors to a select group of trusted parties.
Only about 650 organizations worldwide, including verified government agencies, healthcare providers, and critical infrastructure operators, are eligible to join this program. Interestingly, these partners are provided with a special build that has cybersecurity guardrails completely removed. This is necessary to simulate actual hacking attacks and conduct precise reverse-engineering analysis. The security holes identified this way are then automatically patched in real-time by the 'CodeMender' agent.
You might worry that removing the safety guardrails is too dangerous. To address this, Google has introduced a clever double-monitoring system. This technology monitors the AI's internal 'chain of thought' in real-time while it generates responses. If the analysis shows signs of deviating in a dangerous direction, the system detects this and immediately terminates the process.
Is It a Jack-of-All-Trades? Limitations in Conversational Tasks
While Gemini 4 Argon excels at long-term code analysis and reasoning, it is surprisingly weak at manipulating computer environments in real-time. In the realm of conversational agents that need to watch a screen, move a mouse, or enter commands into a terminal in real-time to solve problems, its responsiveness is somewhat lower compared to competing models.
These weaknesses are clearly visible in performance evaluations. In TerminalBench 4.0, a benchmark test that measures terminal environment analysis capabilities, Argon only reached a 70.6% success rate. This is noticeably disappointing compared to the 70.6% achieved by Anthropic's recently released Claude Sonnet 3.5 or the 66.4% from Claude Opus 3.5.
In OSWorld 2.0, which evaluates the ability to directly control an operating system, Argon recorded 69.2%, falling short of the 72.6% benchmark set by OpenAI's latest model, GPT-6 Astra. While its persistence in deep-diving into complex and long code is excellent, its agility in opening a terminal window and responding quickly in real-time does not yet reach the level of competitors' agents.
Developer Frustration and the Trends Behind the Price Tag
Is it just out of reach? As soon as the overwhelming specs of Gemini 4 Argon were revealed, the global developer community, including Hacker News, was filled with voices of disappointment. Although a powerful model capable of 1 million token output has arrived, ordinary developers are not even getting the chance to try it. Some have expressed deep concerns that as technology advances, the structure where only large corporations and government agencies with capital and power monopolize high-performance models is becoming cemented.
Even with this limited distribution, Google has released a specific API price list for the future full launch. The base service fee is set at $4 per 1 million input tokens and $20 per output token. Fortunately, a promotional period will apply until the end of 2026, offering it at half price: $2 for input and $10 for output. By utilizing the Context Caching feature, which reuses frequently accessed data, users can receive up to a 95% discount on input costs.
However, no matter how attractive the conditions are, it is meaningless if you cannot use it right now. Currently, not only are regular API customers waiting, but even monthly Google AI Ultra subscribers remain on a long waiting list. Many are looking on with disappointment, wondering when the barriers to high-performance AI, locked away for safety, will be opened to ordinary builders.
Restricted High-Performance AI: Will This Become the New Standard?
The distribution method for Gemini 4 Argon could be an important milestone showing how high-performance AI technology will be released to the world in the future. Rather than unconditionally opening technology with high risks of misuse to everyone, a method that restrictively opens access only to verified partners is likely to become a realistic alternative for safe technology supply.
We will be watching with interest to see what results an AI with the overwhelming capability of 1 million token output can achieve on the front lines of security, and whether these tightly closed gates will someday be safely opened for regular developers.