Google didn't announce Gemini Omni. Users found it on their own. Reddit users posted screenshots of a revised Gemini interface exposing a new model card that read: "Create with Gemini Omni: meet our new video model, remix your videos, edit directly in chat, try templates, and more." That's a public-facing UI string, not a dev flag buried in an APK.
A Gemini Omni model has been spotted inside Google's own Gemini video generation interface, fueling speculation that Google is about to launch a single unified model capable of handling text, images, and video in one system. The leak comes just days before Google I/O 2026, scheduled for May 19–20, where the company is expected to make major AI announcements.
As someone who covers this beat closely, this one is worth paying attention to. Not because of the leak itself, but because of what the outputs are already showing.
What Is Gemini Omni
While we've known about Google's Veo model for a while, according to a report from 9to5Google, a brand-new iteration called Gemini Omni is already starting to show up for some users in the wild. At least one Gemini user was prompted to "Create with Gemini Omni," with Google describing the new video generation model as a way to remix your videos, edit directly in chat, and try out pre-made templates.
Toucan is Google's internal codename for the current Veo-3.1-powered video generation pathway inside Gemini. The Omni UI string appeared next to Toucan references, suggesting it may be a replacement or successor.
The name itself carries architectural weight. The most ambitious read is a single Gemini model that handles image generation, video generation, and possibly audio in the same system, the way GPT-4o is positioned for text-image-audio. If true, Gemini would be the first top-tier omni-model with video output.
How the Leak Surfaced
On May 2, 2026, an X user named @Thomas16937378 discovered a UI string inside Google's Gemini video generation tab that read: "Start with an idea or try a template. Powered by Omni." TestingCatalog, a reliable tracker of Google AI leaks, quickly picked up the finding and published a report that spread across the AI community within hours.
A freshly created profile in Gemini's video tab surfaced the "Powered by Omni" line, suggesting the feature is in late-stage testing. This is not a developer build or an APK teardown — it appeared in the live interface.
Two details make it more than noise: the string is visible to users, not just buried in source code or feature flags. UI copy that mentions a brand name typically reaches that state only when the team is preparing for a public release.
Early Output Quality
The early demos are genuinely impressive for a pre-announcement build. Early feedback suggests Gemini Omni may already outperform Veo in several areas. One user praised the model's prompt adherence, smoother camera angle transitions, stronger scene coherence, and significantly improved voice generation quality.
One of the leaked Omni demos specifically used a prompt for two men eating spaghetti at an upscale restaurant, and the results were incredibly realistic. Another demo featured a professor writing out a mathematical proof for trigonometric identities on a traditional chalkboard while explaining the steps. While there are still some obvious AI "tells" in the final output, the model handles text generation and physical movements remarkably well.
Getting math right in video is a hard problem. It requires not just visual coherence but semantic accuracy in the rendered symbols. The fact that an early build is handling it has raised expectations ahead of a formal reveal.
On raw generation fidelity, Omni appears to lag behind ByteDance's Seedance 2, with viewers noting that the cinematic quality is a step behind the current benchmark leader. Where the model stood out was in editing: removing watermarks, swapping objects within clips, and rewriting scenes via chat instructions all worked unusually well for a first public glimpse.
Key Technical Highlights
- Unlike Google's current split approach, where Veo 3.1 handles video generation and separate Gemini-based models handle image generation, Omni could merge these capabilities into a single system that generates text, images, and video natively.
- How "Omni" fits into the broader context of Gemini and Veo isn't entirely clear, but metadata suggests "Omni" is an extension of Veo.
- If the new Omni model extends this to native audio output, not text-to-speech bolted on afterward but audio generated as a first-class output, that would be a meaningful capability jump.
- The model will be available on APIs and is expected to be considered an Agent, similarly to Deep Research on AI Studio.
- Additional leaks suggest Google may release two versions of this model.
The Usage Cost Problem
Generating hyper-realistic video in a chat window takes an immense amount of computing power, and Google is apparently very aware of that. The user who spotted these Omni prompts also noticed a new "usage" tab on their account. Generating just those two video prompts managed to chew through 86% of their daily usage on an AI Pro plan.
Google's intentions to add more explicit usage limits for Gemini have recently been spotted, and a resource-heavy model like Omni is likely the exact reason why.
This is a real constraint for any consumer rollout. If two short clips eat most of a daily quota, the product experience will frustrate users unless Google either optimizes inference costs significantly or tiers access aggressively.
What This Means for the Video AI Race
Google's Veo 3.1 is already considered one of the strongest video generation models on the market, capable of 4K output with natively generated audio. But the leaderboard has remained competitive, with Runway Gen-4.5 having previously edged out Veo 3 on Artificial Analysis benchmarks, and ByteDance's Seedance 2.0 continuing to rank highly on public evaluations.
Current state-of-the-art video models — Veo 3.1, Seedance 2.0, Kling 3.0 — are all specialized video generators. They do not also handle image creation or text reasoning natively. A true omni-model would be structurally different from anything currently available.
If Gemini Omni turns out to be a genuine omni-model, the creator economy implications are significant. The current workflow — storyboard in one tool, still frames in another, video in a third — collapses into a single prompt.
Google I/O 2026 is scheduled for May 19–20. Based on the leaks, we might see a Gemini Omni official announcement, new Gemini 3.2 or 3.5 versions focused on speed and efficiency, and new memory features codenamed "Teamfood" for improved long-term context.
Final Thoughts
The part of this leak that I keep coming back to is the editing capability. Generation quality is a spec race anyone can run. But removing watermarks, swapping objects within clips, and rewriting scenes via chat instructions inside a conversation window is a different kind of product. That's a creative tool, not just a generator.
The compute cost issue is the open question worth watching. If Omni burns through credits at the rate early testers are seeing, Google will need to either compress the model significantly or build a tiered system where casual users get capped access and power users pay more. Neither is trivial to execute well.
Treat all of this as speculative until Google says it on stage. UI strings have shipped without product launches before. But the timing, the live interface placement, and the output quality all point toward a real announcement on May 19. What do you think — is Gemini Omni the video model that finally consolidates Google's fragmented AI stack? Drop your thoughts in the comments.
Frequently Asked Questions
5 questions
1What is Gemini Omni?
Gemini Omni is a leaked Google video generation model spotted in Gemini's UI ahead of Google IO 2026. Evidence suggests it could be the first top-tier omni-model with native video output, potentially replacing Veo 3.1 and unifying image, video, and text generation under one Gemini system.
2How was Gemini Omni discovered?
On May 2, 2026, an X user discovered a UI string in Gemini's video generation tab that read: "Start with an idea or try a template. Powered by Omni." The finding was quickly reported by TestingCatalog, a reliable source for Google AI leaks.
3How does Gemini Omni differ from Veo 3.1?
Gemini already generates videos through its integration with Veo 3.1. The question Omni raises is whether Google is moving from a split-model strategy — Veo for video, separate models for images, Gemini for text — to a unified model that handles all modalities in one system.
4Does Gemini Omni have any usage limits?
The user who spotted the Omni prompts also noticed a new "usage" tab on their account. Generating just two video prompts managed to chew through 86% of their daily usage on an AI Pro plan.
5When will Google officially announce Gemini Omni?
Google I/O 2026 runs May 19–20, 2026. Gemini and AI updates are confirmed agenda items. A pattern of pre-I/O UI leaks surfacing a fresh public name is consistent with a keynote-stage reveal.






