Level 1 — Absolute Beginner
Google makes a new computer brain. Its name is Gemini 4 Argon. It can read, write and help with hard work.
People give tests to these computer brains. Argon wins most of the tests. It beats the brains from two other big companies.
But there is a rule. Most people cannot use it now. Only some safety workers and the government can use it first.
Later, more people can use it. Companies will pay money for it. The price starts low and then goes up.
- computer
- a machine that stores and works with information
- test
- a set of questions used to see how good something is
- win
- to be the best in a game or test
- company
- a business that makes or sells things
- rule
- something you must or must not do
- government
- the group of people who run a country
- pay
- to give money for something
- price
- the money you must give to buy something
Level 2 — Elementary
Google announced a new artificial intelligence model called Gemini 4 Argon on 30 September 2026. The company says it is its most capable system yet, and the published results support that claim.
Argon leads or ties on 13 of the 18 benchmarks Google disclosed. On a legal reasoning test built by the company Harvey, Argon scored 19.6 percent, against 5.4 percent for OpenAI's GPT-6 Astra and 3.8 percent for Anthropic's Claude Opus 5.5. On a software engineering test called DeepSWE it scored 77.9 percent, ahead of Claude at 74.2 and Astra at 74.1.
It does not win everything. Astra leads on a test called FrontierSWE v2 with 65.5 percent against Argon's 55 percent, and Claude Opus 5.5 leads on Terminal-bench 4.0 with 66.4 percent against 57.4 percent.
The unusual part is who gets to use it. Argon goes first to trusted cybersecurity defenders through a Google programme called Fairwind, and to the US government, before reaching ordinary API customers and Google AI Ultra subscribers. Introductory pricing is 2 dollars per million input tokens and 10 dollars per million output tokens, rising later to 4 and 20 dollars.
- artificial intelligence
- computer systems that perform tasks normally needing human thought
- model
- a trained computer system that produces answers from input
- benchmark
- a standard test used to compare performance
- disclose
- to make information public
- reasoning
- the process of thinking through a problem logically
- cybersecurity
- protecting computers and networks from attack
- subscriber
- someone who pays regularly to use a service
- introductory price
- a lower price offered when something first goes on sale
Level 3 — Intermediate
Google unveiled Gemini 4 Argon on 30 September 2026, and on the published evidence it has retaken the performance lead in frontier artificial intelligence. Argon leads or ties on 13 of the 18 benchmarks Google chose to disclose, a tally that places it ahead of OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5 across most of the categories enterprises care about.
The margins vary enormously. On Harvey's legal agent benchmark the gap is almost comic: Argon scored 19.6 percent against 5.4 for Astra and 3.8 for Claude, which says as much about how brutally hard that test is as about Google's progress. On DeepSWE version 1.1, a software engineering evaluation, Argon took 77.9 percent against 74.2 for Claude and 74.1 for Astra, a far narrower win. On AutomationBench it managed 51.3 percent against 42.5 and 41.4.
Rivals retain their own territory. Astra holds FrontierSWE v2 by a wide margin, 65.5 percent to Argon's 55, and Claude Opus 5.5 leads Terminal-bench 4.0 at 66.4 percent against 57.4. Benchmark leadership in 2026 is therefore a patchwork rather than a crown, and buyers increasingly select models by task rather than by brand.
The release strategy is the genuinely novel element. Rather than a broad consumer launch, Google is routing Argon first to vetted cyber defenders through its Fairwind Programme, and granting the US government pre release access, before opening it to API customers and Google AI Ultra subscribers. The company intends to supply trusted defenders with a version stripped of its cyber guardrails, betting that the model's vulnerability remediation skills do more good in defensive hands than harm in open ones. Introductory pricing is 2 dollars per million input tokens and 10 dollars per million output, a fifth of Astra's rate, rising afterwards to 4 and 20.
- frontier model
- the most advanced class of AI system currently built
- tally
- a running count or total
- enterprise
- a large business organisation
- margin
- the amount by which one result exceeds another
- evaluation
- a structured assessment of how well something performs
- patchwork
- something made of many different pieces rather than one whole
- vetted
- checked and approved in advance
- guardrail
- a built in limit that stops a system doing something unsafe
Level 4 — Advanced
Google's announcement of Gemini 4 Argon on 30 September 2026 reclaims a lead the company has surrendered and recovered several times in three years, and the published scorecard is unusually lopsided in its favour. Argon leads or ties 13 of 18 disclosed benchmarks. The headline result, 19.6 percent on Harvey's legal agent evaluation against 5.4 percent for OpenAI's GPT-6 Astra and 3.8 percent for Anthropic's Claude Opus 5.5, looks less like an incremental gain than a discontinuity, though the absolute numbers are a reminder that the task remains largely unsolved by every system tested.
Elsewhere the contest is ordinary. DeepSWE v1.1 separates Argon's 77.9 percent from Claude's 74.2 and Astra's 74.1 by a margin well inside the range where evaluation design and prompt formatting can move results. AutomationBench shows a wider spread at 51.3 against 42.5 and 41.4. Against that, Astra retains FrontierSWE v2 at 65.5 percent to 55, and Claude Opus 5.5 holds Terminal-bench 4.0 at 66.4 to 57.4. The sensible reading is that no laboratory now holds a general lead, only a shifting portfolio of task specific ones, and procurement decisions have begun to reflect that.
The distribution choice is more interesting than the scores. Google is withholding general availability and routing Argon first through its Fairwind Programme to vetted cyber defenders, with pre release access for the United States government, ahead of any broader rollout to API customers and Google AI Ultra subscribers. More striking still, the company intends to supply trusted defenders with a build whose cyber guardrails have been removed, on the reasoning that a model competent at vulnerability remediation is more valuable inside defensive institutions than it is dangerous there.
That is a defensible bet and an unavoidably political one. It presumes that the set of trustworthy recipients can be defined and policed, that a capability released to defenders does not leak to the people they defend against, and that a private company is the appropriate body to draw that boundary. Commercially, the pricing suggests Google expects to compete on cost as well as capability: 2 dollars per million input tokens and 10 per million output at introduction, roughly a fifth of Astra's rate and half of Claude Opus 5.5, rising to 4 and 20 once the promotional window closes. The context window now extends to a million output tokens, up from 64,000. Under DeepMind's new chief, Koray Kavukcuoglu, Google appears to have concluded that frontier capability is worth more as a security asset than as a consumer product, at least for a quarter.
- lopsided
- heavily unbalanced towards one side
- discontinuity
- a sudden break from a previous pattern rather than gradual change
- incremental
- happening in small steps rather than one leap
- procurement
- the process by which an organisation buys goods or services
- general availability