Guides Google published sep 30, 2026

Gemini 4 Argon vs Gemini Flash and Pro: Which Model Should You Use?

beginner

The Rundown

Updated September 30, 2026

Google announced Gemini 4 Argon on September 30, 2026.

Argon targets work that takes many steps: complex code changes, research across documents, and business tasks that need sustained reasoning. Google DeepMind also emphasizes its ability to find and fix software security flaws.

Our starting recommendation: use Gemini 3.8 Flash for substantial everyday work, try 3.5 Flash-Lite for simple repeated tasks, and compare Argon on harder jobs when you gain access. Keep a working Pro setup until a comparison gives you a reason to change.

This guide covers access, performance, pricing, and a short decision process. You do not need to build an app to use it.

Google's Gemini 4 Argon launch artwork
Google's Gemini 4 Argon launch artwork

Image credit: Google.

Can you use Gemini 4 Argon now?

Access starts with selected testers. Google is giving a set of trusted cybersecurity partners access through its Fairwind Program. The program vets applicants and restricts how they can use and share access. It is not a general consumer waitlist.

Google plans to expand access, starting with paid API customers and Google AI Ultra subscribers. It has not given a firm date for that wider release. Check availability before upgrading your plan for Argon; a paid plan does not establish access today.

At this check, Google's public Gemini API model catalog still lists the Gemini 3 models below and has no Argon entry. We therefore have no public model ID or general setup steps to verify yet.

Use the comparison to choose something available now and decide where Argon deserves a test later.

Gemini 4 Argon, Flash, Flash-Lite, and Pro compared

This table uses exact model versions from Google's developer documentation. The task examples are our suggested starting points.

ModelCurrent statusWhere we would start
Gemini 4 ArgonSelected early access; wider release pendingDifficult code changes, long research tasks, and work that repeatedly defeats your current model.
Gemini 3.8 FlashStable API modelReports, document analysis, coding, and other substantial daily work.
Gemini 3.5 Flash-LiteStable API modelSorting requests, extracting fields, and processing many short documents.
Gemini 3.1 Pro PreviewPublic API previewExisting Pro workflows and comparison runs on difficult reasoning tasks.

Gemini 4 Argon vs Gemini 3.8 Flash

Google positions Argon for sustained reasoning across complex work. Gemini 3.8 Flash already supports substantial coding and business tasks, so a difficult job can still be a reasonable Flash test.

Our rule: start with Flash when you can define and check the result. Save the failures worth testing with Argon, such as an overlooked condition in a project brief or a bug that survives a proposed fix.

Gemini 4 Argon vs Gemini 3.1 Pro

Google still lists 3.1 Pro Preview for complex reasoning and coding. Its preview status matters when deciding what to use in an established app.

Keep examples from your current Pro setup. Compare whether another model catches more errors, follows the required format, or reduces editing. The model name alone gives you too little evidence to justify changing a working process.

Where Flash-Lite fits

Google designs 3.5 Flash-Lite for fast, high-volume work, including extraction and document processing.

Try it on a narrow task with clear rules. For example, extract company names and dates from approved notes, or sort support requests into five fixed categories. Send ambiguous cases for review. A smaller task makes it easier to judge whether the result meets your needs.

A much larger output limit

Google announced a maximum of 1 million output tokens, up from the previous 64K limit. A token is a small unit of information the model processes.

Keep two limits separate when reading a model specification:

LimitWhat it tells you
Input limitHow much source material the model can accept in a request.
Output limitThe capacity available for the model's generated output.

For comparison, 3.8 Flash, 3.5 Flash-Lite, and 3.1 Pro Preview each list 1,048,576 input tokens and 65,536 output tokens in their API specifications.

The Argon launch post does not state a separate public input limit. Avoid treating its output announcement as proof of a larger input window.

For a weekly report, you may still want 500 words. More capacity is useful only when it helps the model complete the task correctly. Keep your requested answer length tied to what someone needs to read.

Stronger results on a published work benchmark

The Vals Index v2.1 combines tests covering finance, coding, legal, and tax work. Its live table showed these results when we checked it on September 30:

Google model in the Vals tableVals Index score
Gemini 4 Argon68.90%
Gemini 3.8 Flash54.83%
Gemini 3.7 Flash51.27%
Gemini 3.1 Pro Preview, labeled 02/2633.44%

Argon leads these Google models on this test. The result supports testing it on complex professional work. It does not establish an accuracy rate for your own documents or an equal-cost comparison: the runs use different token totals and durations.

Use the benchmark to choose candidates. Then check them on work where you know what a correct answer requires.

Google's broader benchmark comparison

The chart below compares Argon with other providers' frontier models, not Flash and Pro. The supplied chart matches Google DeepMind's published results. Argon leads some tests and trails others, so it is not a universal ranking of model quality.

Google's benchmark matrix comparing Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5
Google's benchmark matrix comparing Gemini 4 Argon, GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5

Chart credit: Google DeepMind. Evaluation details are linked in Google's methodology page. We have not independently reproduced these results.

Coding and cybersecurity results

Two reported results are especially relevant to Argon's positioning:

Google-reported benchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
DeepSWE v1.1, long-horizon software engineering77.9%74.1%67.4%74.2%
CWE-bench v1, security vulnerability remediation68.0%68.0%58.0%67.0%
Google's DeepSWE v1.1 chart: Argon 77.9%, GPT-6 Astra 74.1%, Claude Fable 5.1 67.4%, and Claude Opus 5.5 74.2%
Google's DeepSWE v1.1 chart: Argon 77.9%, GPT-6 Astra 74.1%, Claude Fable 5.1 67.4%, and Claude Opus 5.5 74.2%

Chart credit: Google. Higher is better on this specific software-engineering evaluation.

Google's CWE-bench v1 leaderboard: Gemini 4 Argon, Grok 4.7, and GPT-6 Astra each score 68%
Google's CWE-bench v1 leaderboard: Gemini 4 Argon, Grok 4.7, and GPT-6 Astra each score 68%

Chart credit: Google DeepMind. The chart reports Pass@1, with ties broken by Pass@4, and identifies each model's agent harness. Treat these as model-and-harness results, not proof that Argon alone makes a workflow secure.

Gemini 4 Argon pricing vs Flash and Pro

The table shows US dollars per million API tokens. The final column illustrates 20,000 uncached input tokens plus 5,000 billed output tokens. It holds usage equal across models. Current-model rates come from Google's API pricing; Argon's rates are announced launch prices, not evidence of general availability.

Model and pricing periodInputOutputExample token charge
Gemini 3.5 Flash-Lite$0.30$2.50$0.0185
Gemini 3.8 Flash, through December 31, 2026$0.75$3.75$0.03375
Gemini 3.1 Pro Preview, prompts up to 200K tokens$2.00$12.00$0.10
Gemini 4 Argon, announced introductory rate$2.00$10.00$0.09
Gemini 4 Argon, after the introductory period$4.00$20.00$0.18

Argon's introductory end date remains unspecified.

Flash's price change has a date: 3.8 Flash moves to $1.50 input and $7.50 output per million tokens on January 1, 2027. Keep that change in any longer-term budget; it is not Argon's introductory deadline.

The Vals leaderboard's run-cost columns use Argon's later $4/$20 rates and Flash's later $1.50/$7.50 rates. They are not priced on the same basis as the current-rate illustration above.

For 3.1 Pro Preview, prompts above 200K input tokens cost $4 input and $18 output per million. The existing Gemini 3 output rates include thinking tokens. Caching, storage, tools, and processing options can change the bill. These figures describe API use; Gemini app plans have separate allowances.

Our advice: budget for the rate that will apply when you use the model. Argon's introductory price makes a trial less costly than its announced later rate. The result still needs to earn the extra cost over a smaller model.

The example amounts are calculations, not measured task costs or live model results.

Learn to compare models with your own work

In Intro to AI Evals, Nate Grahek teaches how to build tests from real inputs, define expected answers, and compare quality, speed, and cost. The course uses spreadsheet and Claude Code examples you can adapt to your own model decisions.

Watch Intro to AI Evals with University Pro →

The course teaches evaluation methods; it does not grant Argon access. Your Google plan and API usage have separate access and billing terms.

Verification: We checked the linked announcement, model specifications, pricing, app help pages, and benchmark tables on September 30, 2026. We verified the price calculations and the supplied charts against Google's published figures. We have not run a live Argon comparison, independently audited the full evaluation methodology, or tested its controls in a signed-in account.

$29 billed monthly · cancel anytime
Go Pro