A few weeks ago, I would have easily said Seedance 2.0 was one of the most capable video models available right now. It is fast, good at hyperrealism, and can handle physics really well. It also has one of the more interesting creative workflows because of its Omni feature, where you can drop different media types into one dashboard and let the system understand how they should be used in the final video.
But now, Google may be preparing something that could challenge that.
A new video model called Gemini Omni appears to have leaked ahead of Google I/O 2026.
The first leak came from TestingCatalog, which spotted a Gemini video generation interface with the phrase, “Start with an idea or try a template. Powered by Omni.” Here’s the screenshot:

Gemini Omni leak. Image by TestingCatalog
The timing also makes this leak more interesting. Google I/O 2026 is happening on May 19 and May 20, and Google is expected to talk heavily about Gemini, AI products, and probably its next wave of media generation tools.
Then the sample videos allegedly generated with Gemini Omni started circulating on X. Here’s one:
As a proof that this video was indeed generated with Gemini, you can check the original source of the video here: https://gemini.google.com/share/7d5dc678c80a
The actual prompt used for this video:
Prompt : A professor writes out a mathematical proof for trigonometric identities on a traditional chalkboard, explaining the step he is currently on in the equation.
That is not a simple prompt. A professor writing on a board involves body movement, hand movement, chalk movement, readable text, scene consistency, and a very specific educational context. A lot of other video easily fails when you ask for meaningful actions with text and hands in the same scene.
The Spaghetti Test
Another leaked demo reportedly showed two men eating spaghetti in a restaurant. If you remember the early days of AI video, you probably remember the horrifying viral clip of Will Smith eating spaghetti. It became the unofficial benchmark for how bad AI video could be.
That is why I think this leak deserves attention. Not because every leaked AI model should be hyped, but because Omni could represent a more serious direction for Google’s video strategy.
If the leak is accurate, this might not be just another Veo update. It could be Google’s attempt to make video generation feel more native inside Gemini, with chat-based editing, templates, remixing, and possibly a more multimodal creative workflow.
What we know about Omni so far
Right now, Google has not officially announced Gemini Omni, so we need to be careful. There is a difference between confirmed information and speculation based on leaks.
What we actually have so far is a mix of UI screenshots, early demo clips, and reports from people tracking Gemini’s upcoming features.
Another detail is that Omni was reportedly seen near Toucan, which has been described as Google’s current video generation pathway linked to Veo 3.1. This is where things become more interesting. If Omni is appearing beside the existing video pipeline, then it could mean Google is preparing a new model, a new interface, or a replacement for the current Veo-powered experience.
The name Omni does not confirm much yet, but it hints at a broader direction.
Based on the leaks, there are two likely possibilities. Gemini Omni could simply be a new Gemini interface built on top of Google’s existing Veo video system. If that is the case, Omni may be less about a new model and more about a better creative workflow inside Gemini, with templates, remixing, and chat-based editing.
There was also a prompt spotted that said “Create with Gemini Omni.” According to Chrome Unboxed and other reports, the feature appeared to describe a system that can remix videos, edit directly in chat, try templates, and perform other creative tasks.
The more interesting possibility is that Omni is a new multimodal video model. As the name suggests, it could understand more than text prompts. It may be designed to work with images, videos, references, templates, and possibly audio in one flow. That would make it closer to the Omni feature in Seedance 2.0, where you can drop different media types into one dashboard and let the system understand how to use them for video generation.
That is the part that makes this leak more interesting to me. Most people do not create videos from a blank prompt. They start with a product photo, a character image, a short script, a previous clip, a brand reference, or a rough idea. If Gemini Omni can understand those inputs directly inside Gemini, then it could become more than another text-to-video model.
Of course, this is still speculation. The leak points in that direction, but it does not prove the full capability yet.
When will it be announced?
The most likely announcement window is Google I/O 2026, which is scheduled for May 19 to May 20.
I do not expect Google to make this fully open to everyone right away. Video generation is expensive, and if the reports about high usage cost are accurate, Omni may launch first as a limited preview or as part of a paid Gemini plan.
Google will obviously show the best clips. What I want to see is the actual workflow.
- Can users upload reference images?
- Can they edit generated videos inside chat?
- Can Omni preserve characters across scenes?
- Can it handle readable text better than current models?
- Can it understand multiple media inputs at once?
And most importantly, can regular users generate enough videos without burning through their limits too fast?
Final Thoughts
Gemini Omni is still unofficial, but the leak gives us a good idea of where Google may be heading. The model name appeared inside Gemini’s video creation interface, sample videos started circulating online, and reports suggest that Omni may support more than basic text-to-video generation.
It could include templates, remixing, and direct chat-based editing. If that is true, then Omni may not simply be a Veo replacement. It could be Google’s attempt to build a more complete creative video system inside Gemini.
Personally, I do not want another model that only produces better-looking videos. We already have Seedance 2.0. What I want is a video model that understands context, references, media inputs, and revisions.
If Google is building a model that can understand text, image, video, and maybe audio in one creative workflow, then the next few months in AI video could be very interesting.
What do you think about Gemini Omni? Are you excited to see Google’s next video model, or do you think Seedance 2.0 will stay ahead for now?
Let me know what you think in the comments.
Sources
- Seedance 2.0generativeai.pub
- TestingCatalogtestingcatalog.com
- https://x.com/chetaslua/status/2053824398503678108x.com
- https://gemini.google.com/share/7d5dc678c80agemini.google.com
- https://vimeo.com/1191453074?fl=pl&fe=vlvimeo.com
- Chrome Unboxedchromeunboxed.com
