Keep going
Keep learning with Rundown Pro
Pro adds the full on-demand course library, certifications, group office hours, and eligible member perks.
$29/mo, cancel anytime
Guides Google published sep 30, 2026
Updated September 30, 2026
Google announced Gemini 4 Argon on September 30, 2026.
Argon targets work that takes many steps: complex code changes, research across documents, and business tasks that need sustained reasoning. Google DeepMind also emphasizes its ability to find and fix software security flaws.
Our starting recommendation: use Gemini 3.8 Flash for substantial everyday work, try 3.5 Flash-Lite for simple repeated tasks, and compare Argon on harder jobs when you gain access. Keep a working Pro setup until a comparison gives you a reason to change.
This guide covers access, performance, pricing, and a short decision process. You do not need to build an app to use it.
Image credit: Google.
Access starts with selected testers. Google is giving a set of trusted cybersecurity partners access through its Fairwind Program. The program vets applicants and restricts how they can use and share access. It is not a general consumer waitlist.
Google plans to expand access, starting with paid API customers and Google AI Ultra subscribers. It has not given a firm date for that wider release. Check availability before upgrading your plan for Argon; a paid plan does not establish access today.
At this check, Google's public Gemini API model catalog still lists the Gemini 3 models below and has no Argon entry. We therefore have no public model ID or general setup steps to verify yet.
Use the comparison to choose something available now and decide where Argon deserves a test later.
This table uses exact model versions from Google's developer documentation. The task examples are our suggested starting points.
| Model | Current status | Where we would start |
|---|---|---|
| Gemini 4 Argon | Selected early access; wider release pending | Difficult code changes, long research tasks, and work that repeatedly defeats your current model. |
| Gemini 3.8 Flash | Stable API model | Reports, document analysis, coding, and other substantial daily work. |
| Gemini 3.5 Flash-Lite | Stable API model | Sorting requests, extracting fields, and processing many short documents. |
| Gemini 3.1 Pro Preview | Public API preview | Existing Pro workflows and comparison runs on difficult reasoning tasks. |
Google positions Argon for sustained reasoning across complex work. Gemini 3.8 Flash already supports substantial coding and business tasks, so a difficult job can still be a reasonable Flash test.
Our rule: start with Flash when you can define and check the result. Save the failures worth testing with Argon, such as an overlooked condition in a project brief or a bug that survives a proposed fix.
Google still lists 3.1 Pro Preview for complex reasoning and coding. Its preview status matters when deciding what to use in an established app.
Keep examples from your current Pro setup. Compare whether another model catches more errors, follows the required format, or reduces editing. The model name alone gives you too little evidence to justify changing a working process.
Google designs 3.5 Flash-Lite for fast, high-volume work, including extraction and document processing.
Try it on a narrow task with clear rules. For example, extract company names and dates from approved notes, or sort support requests into five fixed categories. Send ambiguous cases for review. A smaller task makes it easier to judge whether the result meets your needs.
Google announced a maximum of 1 million output tokens, up from the previous 64K limit. A token is a small unit of information the model processes.
Keep two limits separate when reading a model specification:
| Limit | What it tells you |
|---|---|
| Input limit | How much source material the model can accept in a request. |
| Output limit | The capacity available for the model's generated output. |
For comparison, 3.8 Flash, 3.5 Flash-Lite, and 3.1 Pro Preview each list 1,048,576 input tokens and 65,536 output tokens in their API specifications.
The Argon launch post does not state a separate public input limit. Avoid treating its output announcement as proof of a larger input window.
For a weekly report, you may still want 500 words. More capacity is useful only when it helps the model complete the task correctly. Keep your requested answer length tied to what someone needs to read.
The Vals Index v2.1 combines tests covering finance, coding, legal, and tax work. Its live table showed these results when we checked it on September 30:
| Google model in the Vals table | Vals Index score |
|---|---|
| Gemini 4 Argon | 68.90% |
| Gemini 3.8 Flash | 54.83% |
| Gemini 3.7 Flash | 51.27% |
| Gemini 3.1 Pro Preview, labeled 02/26 | 33.44% |
Argon leads these Google models on this test. The result supports testing it on complex professional work. It does not establish an accuracy rate for your own documents or an equal-cost comparison: the runs use different token totals and durations.
Use the benchmark to choose candidates. Then check them on work where you know what a correct answer requires.
The chart below compares Argon with other providers' frontier models, not Flash and Pro. The supplied chart matches Google DeepMind's published results. Argon leads some tests and trails others, so it is not a universal ranking of model quality.
Chart credit: Google DeepMind. Evaluation details are linked in Google's methodology page. We have not independently reproduced these results.
Two reported results are especially relevant to Argon's positioning:
| Google-reported benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|
| DeepSWE v1.1, long-horizon software engineering | 77.9% | 74.1% | 67.4% | 74.2% |
| CWE-bench v1, security vulnerability remediation | 68.0% | 68.0% | 58.0% | 67.0% |
Chart credit: Google. Higher is better on this specific software-engineering evaluation.
Chart credit: Google DeepMind. The chart reports Pass@1, with ties broken by Pass@4, and identifies each model's agent harness. Treat these as model-and-harness results, not proof that Argon alone makes a workflow secure.
The table shows US dollars per million API tokens. The final column illustrates 20,000 uncached input tokens plus 5,000 billed output tokens. It holds usage equal across models. Current-model rates come from Google's API pricing; Argon's rates are announced launch prices, not evidence of general availability.
| Model and pricing period | Input | Output | Example token charge |
|---|---|---|---|
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | $0.0185 |
| Gemini 3.8 Flash, through December 31, 2026 | $0.75 | $3.75 | $0.03375 |
| Gemini 3.1 Pro Preview, prompts up to 200K tokens | $2.00 | $12.00 | $0.10 |
| Gemini 4 Argon, announced introductory rate | $2.00 | $10.00 | $0.09 |
| Gemini 4 Argon, after the introductory period | $4.00 | $20.00 | $0.18 |
Argon's introductory end date remains unspecified.
Flash's price change has a date: 3.8 Flash moves to $1.50 input and $7.50 output per million tokens on January 1, 2027. Keep that change in any longer-term budget; it is not Argon's introductory deadline.
The Vals leaderboard's run-cost columns use Argon's later $4/$20 rates and Flash's later $1.50/$7.50 rates. They are not priced on the same basis as the current-rate illustration above.
For 3.1 Pro Preview, prompts above 200K input tokens cost $4 input and $18 output per million. The existing Gemini 3 output rates include thinking tokens. Caching, storage, tools, and processing options can change the bill. These figures describe API use; Gemini app plans have separate allowances.
Our advice: budget for the rate that will apply when you use the model. Argon's introductory price makes a trial less costly than its announced later rate. The result still needs to earn the extra cost over a smaller model.
The example amounts are calculations, not measured task costs or live model results.
In Intro to AI Evals, Nate Grahek teaches how to build tests from real inputs, define expected answers, and compare quality, speed, and cost. The course uses spreadsheet and Claude Code examples you can adapt to your own model decisions.
Watch Intro to AI Evals with University Pro →
The course teaches evaluation methods; it does not grant Argon access. Your Google plan and API usage have separate access and billing terms.
Verification: We checked the linked announcement, model specifications, pricing, app help pages, and benchmark tables on September 30, 2026. We verified the price calculations and the supplied charts against Google's published figures. We have not run a live Argon comparison, independently audited the full evaluation methodology, or tested its controls in a signed-in account.
Choose a task you do often enough to test, then use this decision table.
| Your task | Our starting choice | What would justify changing it? |
|---|---|---|
| Extract names, dates, or labels from many short items. | 3.5 Flash-Lite | It misses required fields or invents missing values. |
| Turn notes and results into a report. | 3.8 Flash | It repeatedly misses a conflict, calculation, or required section. |
| Make and test a code change. | 3.8 Flash, alongside your current model | The patch fails the tests or changes unrelated behavior. |
| Maintain a workflow already tested with Pro. | Your current Pro setup | Another model passes the same checks with less review or cost. |
| Resolve a hard, multi-step problem that current models keep missing. | Compare Argon when available | It fixes the observed failure at an acceptable cost. |
These are editorial recommendations based on the roles Google documents for Flash, Flash-Lite, Pro, and Argon. They are starting points for a test, rather than a measured ranking for these tasks.
In the Gemini app, Google's help page describes Flash-Lite, Flash, and Pro choices. It does not map those labels to every API version in this guide. Check what your account offers and record the selected model and thinking level.
For Gemini 3.8 Flash in the API, Google supports low, medium, and high thinking levels, with medium as the default. More thinking can increase time and token use. Try a higher setting when you can name the error it needs to fix.
Deep Think is another option for eligible app users. Google documents it under Pro → Thinking Level → Deep Think for adults with Google AI Ultra or an eligible Ultra for Business license. It is experimental and can take several minutes. Treat it as a separate setting to test.
Keep the prompt and tool access consistent when comparing runs. Record the thinking settings; matching names across models do not establish equal amounts of computation.
Choose a few examples with known answers. Include a normal case, an incomplete brief, and a case your current model has already failed. Use fictional or approved material; do not upload private work without permission.
| Example | Checks to make before accepting the answer |
|---|---|
| Marketing report | Calculations match the source; uncertain causes stay uncertain; dates and reporting periods match. |
| Meeting follow-up | Every owner and deadline has support in the notes; suggestions stay separate from agreed tasks. |
| Code change | The stated bug is fixed; relevant tests pass; unrelated behavior stays intact. |
Save the first answer from each model before asking for revisions. Record mistakes, time spent editing, and the final cost where available. Give every model the same chance to revise, or compare first attempts only.
Choose the model that meets your checks with the least total work and expense. A longer answer or a more confident tone should carry little weight on its own.
Keep going
Pro adds the full on-demand course library, certifications, group office hours, and eligible member perks.
$29/mo, cancel anytime
Google's public API catalog still lists Gemini 3.8 Flash, 3.5 Flash-Lite, and 3.1 Pro Preview. Keep existing integrations until you have confirmed a supported upgrade path and tested it.
Google presents Argon as a model. Its help center presents Deep Think as an experimental reasoning option selected under Pro in the Gemini app. Keep their access and settings separate.
Check Google's specialized models. The public catalog lists Nano Banana models for image creation and editing, and Gemini 3.8 Live for real-time voice. Choose a model that supports the output your task needs.
Start with a few hard tasks you can judge. Keep your existing setup until the new model improves results, time, or cost. For an app, also check its supported tools, output format, limits, and failure handling before changing production traffic.