Technical specifications of Gemini 4 Argon
| Specification | Gemini 4 Argon |
|---|---|
| Provider | Google DeepMind |
| Model family | Gemini 4 |
| Model ID | gemini-4-argon |
| Model type | Frontier multimodal reasoning model |
| Primary focus | Complex workflows, agentic coding, enterprise knowledge work, cybersecurity |
| Multimodal capability | Google describes multimodal understanding; the public model page does not yet provide a complete API modality schema |
| Maximum output | Up to 1,000,000 tokens, according to current third-party API listings; this should not be confused with the input context window |
| Input context window | Not clearly published by Google on the current public model page |
| Public API availability | Limited/pre-release access; Google has not announced a general public release date |
What is Gemini 4 Argon?
Gemini 4 Argon is the flagship model anchoring Google's Gemini 4 generation. Google describes it as designed for frontier performance on complex workloads, including software engineering, enterprise knowledge work, and cybersecurity defense. Its published evaluation set covers knowledge work, agentic coding, science and mathematics, computer use, multimodal understanding, long-context tasks, and cybersecurity.
Argon's positioning is notably workflow-oriented rather than limited to conversational question answering. Google highlights the ability to sustain long, multi-step tasks and to operate across complex software-engineering and enterprise workflows.
Main features of Gemini 4 Argon
- Complex workflow reasoning: Designed for long, multi-step tasks that require sustained reasoning rather than a single short response.
- Agentic coding: Google reports strong results on DeepSWE v1.1, Vibe Code Bench, and other coding evaluations.
- Enterprise knowledge work: Published evaluations include finance and legal-agent tasks as well as end-to-end business automation.
- Multimodal understanding: Google reports strong results on Chartography and LVBench, covering multimodal and long-video understanding.
- Long-context performance: GraphWalks results are reported for both up-to-128K and 256K-to-1M-token ranges.
- Cybersecurity defense: Google specifically positions Argon for defensive cybersecurity and reports CWE-bench v1 results for vulnerability remediation.
Benchmark performance of Gemini 4 Argon
Google DeepMind reports the following results on its current Gemini 4 Argon evaluation page. These are vendor-published results and should be compared only within the stated benchmark methodology. [1]
| Benchmark | Category | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
| Vals Index | Knowledge work | 68.9% | 63.1% | 65.8% | 67.0% |
| AutomationBench | Knowledge work | 51.3% | 41.4% | 31.4% | 42.5% |
| Vals Finance Agent v2 | Knowledge work | 65.4% | 53.5% | 58.9% | 58.6% |
| Harvey's Legal Agent Benchmark | Knowledge work | 19.6% | 5.4% | 6.7% | 3.8% |
| DeepSWE v1.1 | Agentic coding | 77.9% | 74.1% | 67.4% | 74.2% |
| FrontierSWE v2 | Agentic coding | 55.0% | 65.5% | 56.3% | 62.3% |
| Vibe Code Bench | Agentic coding | 91.9% | 89.6% | 90.3% | 90.3% |
| Terminal-bench 4.0 | Agentic coding | 57.4% | 58.2% | 57.9% | 66.4% |
| Terminal-Bench Science 0.1 | Science & math | 57.6% | 68.1% | 52.6% | 63.3% |
| LABBench | Science & math | 88.8% | 85.4% | 68.6% | 73.1% |
| RiemannBench | Science & math | 76.0% | 72.0% | 65.6% | 69.6% |
| GraphWalks, up to 128K | Long context | 99.7% | 98.7% | 91.4% | 90.6% |
| GraphWalks, 256K–1M | Long context | 84.2% | 71.8% | 65.0% | 66.8% |
| Chartography | Multimodal understanding | 71.6% | 71.0% | 46.2% | 66.3% |
| LVBench | Multimodal understanding | 91.7% | 87.5% | 79.7% | 83.7% |
| CWE-bench v1 | Cybersecurity | 68.0% | 68.0% | 58.0% | 67.0% |
The results show that Argon's performance varies by task: it leads the cited comparison on several knowledge-work, coding, science, long-context, and multimodal evaluations, while other models score higher on specific coding, science, or computer-use benchmarks. They should not be collapsed into a single overall ranking.
Gemini 4 Argon vs GPT-6 Astra vs Claude Fable 5.1
| Area | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 |
|---|---|---|---|
| Knowledge work | Strong published results across Vals, AutomationBench, finance, and legal-agent evaluations | Strong on the same evaluation set | Strong on several knowledge-work evaluations |
| Agentic coding | 77.9% on DeepSWE v1.1; 91.9% on Vibe Code Bench | 74.1% on DeepSWE v1.1; 89.6% on Vibe Code Bench | 67.4% on DeepSWE v1.1; 90.3% on Vibe Code Bench |
| Long context | 99.7% on GraphWalks up to 128K; 84.2% from 256K–1M | 98.7% and 71.8% respectively | 91.4% and 65.0% respectively |
| Multimodal understanding | 71.6% Chartography; 91.7% LVBench | 71.0%; 87.5% | 46.2%; 79.7% |
| Cybersecurity | 68.0% on CWE-bench v1 | 68.0% | 58.0% |
The comparison is benchmark-specific. For example, GPT-6 Astra scores higher than Argon on FrontierSWE v2 and Terminal-Bench Science 0.1, while Argon scores higher on DeepSWE v1.1, Vibe Code Bench, GraphWalks 256K–1M, and LVBench.
Representative use cases
- Repository-scale software engineering — Long-horizon coding, refactoring, debugging, and agentic development.
- Enterprise knowledge workflows — Legal research, financial analysis, and business-process automation.
- Cybersecurity defense — Vulnerability discovery, validation, and remediation in controlled environments.
- Scientific and mathematical workflows — Tasks reflected by LABBench, RiemannBench, and science-oriented evaluations.
- Long-video and multimodal analysis — Workloads requiring interpretation across visual and temporal information.
- Long-context agents — Applications that need to maintain information across very large task trajectories.