Google employees are reportedly testing a Gemini 4 checkpoint called Carbon inside the company’s internal coding platform, Jetski. According to Business Insider’s October 9 report, which cites documents and screenshots reviewed by the publication, the model became available internally in recent days.
The report includes encouraging employee assessments of Carbon’s coding abilities, including a comparison with Anthropic’s Claude Opus 5.5. Those comments are early impressions, not published benchmark results.
The development is notable because Google’s first announced Gemini 4 model, Argon, was still awaiting a broad public release at the time of the report. Business Insider’s reporting suggests Google is already testing a newer checkpoint behind that rollout. It does not establish that Carbon will become a product: Google declined to comment to the publication, and its eventual name, availability, and relationship to Argon remain unresolved.
Carbon’s Internal Name Does Not Establish Its Release Path
Business Insider describes Carbon as a new version of Gemini 4, possibly an update to Argon, being tested through Jetski. That places the reported activity in an internal coding environment rather than a public launch.
A checkpoint generally refers to a saved state of a model during development. Testing a checkpoint can help assess a candidate’s capabilities, but the label alone does not establish a new commercial model, a different architecture, or a release decision. Here, the useful distinction is between what employees can reportedly access inside Google and what developers can obtain outside it.
The naming history in Business Insider’s report makes that distinction particularly important. According to an internal document reviewed by the publication, Google has tested Gemini 4 models under the names Argon, Barium, and Carbon. The document said that an internal model called Barium-B was selected to become the model publicly known as Argon.
That mapping prevents a straightforward reading of the codenames as a public product sequence. An internal Barium checkpoint can become public-facing Argon; Carbon could conceivably follow a similar naming change. Business Insider could not determine whether Carbon would arrive as another Argon update or as a separate model in the Gemini 4 family.
The report therefore supports a narrower conclusion than “Google is launching Gemini 4 Carbon.” It describes an internal development candidate, not a confirmed addition to Google’s public model lineup.
The Claude Opus 5.5 Comparison Is an Employee Impression
The most attention-grabbing assessment comes from an employee who told Business Insider that Carbon “feels like Opus 5.5” for coding. The same employee cautioned that more testing was needed.
That caveat belongs alongside the comparison. It is an individual’s assessment of an internal model, not evidence that Carbon matches Anthropic’s Claude Opus 5.5 across coding tasks.
Business Insider also reported positive comments in internal messaging channels, including another employee’s description of Carbon as “really good.” These reactions indicate enthusiasm among the employees cited. They do not establish how broadly that view is shared inside Google, or how Carbon would perform for outside developers.
The report provides some context for the enthusiasm. According to Business Insider, early Argon versions were generally well received by staff, but one employee thought they lagged on some coding tasks and compared them with the older Claude Opus 5. Those observations suggest that at least some staff perceive progress between earlier Gemini 4 candidates and Carbon. They cannot quantify that progress.
A meaningful comparison would need to identify the exact model versions, tasks, tools, and evaluation conditions. For agentic coding, an assessment should also distinguish producing a plausible edit from completing a task successfully: changing the right files, passing relevant tests, and avoiding regressions.
None of those details accompanies the employee comparison in the supplied reporting. Nor does the report provide a published independent Carbon benchmark.
Anthropic’s Claude serves as a reference point in the employee feedback, but the evidence supports reporting that comparison, not adopting it as a performance verdict. Carbon may warrant attention from developers following the coding-model race; declaring parity with Opus 5.5 would go further than the available evidence allows.
Argon’s Announcement Provides Context, Not Carbon Specifications
Google announced Gemini 4 Argon on September 30. In its official Argon announcement, the company said the model was rolling out to trusted cyber defenders through its Fairwind Program, with broader availability to follow.
Google did not give a public-release date. It described a phased approach involving early testing, safeguards, and participation in the U.S. government’s voluntary process for pre-release model access. Business Insider’s September 30 coverage likewise distinguished the initial partner access from a future public rollout.
Argon was therefore announced, but not generally available. That is the backdrop for Carbon’s reported internal testing.
Google positioned Argon around complex software engineering, enterprise knowledge work, and cybersecurity defense. The company also reported internal uses involving debugging, codebase migrations, and algorithm design. These are Google’s descriptions of Argon, not independent findings about Carbon.
The same separation applies to performance claims. Google reported a 77.9% result for Argon on DeepSWE v1.1, which it describes as measuring real-world, long-horizon software engineering tasks. That is a company-reported result for a named Argon evaluation. It cannot establish Carbon’s score or validate an employee’s comparison with Claude.

Likewise, Argon’s announced pricing and technical specifications should not be transferred to Carbon. Business Insider’s report does not establish that the internal checkpoint will inherit Argon’s commercial terms or public configuration.
The two sources answer different questions. Google’s announcement explains the model it has chosen to introduce and its initial access plan. Business Insider’s reporting describes additional internal development. Reading them together suggests that the publicly announced model and the latest reported internal candidate are not necessarily the same thing.
Sources
- Business Insider’s October 9 reportbusinessinsider.com
- official Argon announcementblog.google
- September 30 coveragebusinessinsider.com





