xAI’s new Grok Imagine Video 1.5 Lite starts at $0.02 per generated second, but that price applies to 480p output. Moving to 1080p raises the Lite rate to $0.14 per second, and the resulting video is upscaled rather than rendered natively at that resolution.
That distinction matters for developers choosing between Lite and the fuller Grok Imagine Video 1.5 model. Both are now listed in xAI’s Imagine API overview, but the full model adds native 1080p generation, reference-image controls, preset voice references and pinned frames.
The practical choice is between a lower-cost generation tier and a model with more production controls. The documentation establishes those differences; it does not establish how consistently either model produces usable footage or independently verify a quality advantage.
The $0.02 Price Applies Only to 480p
The Grok Imagine Video 1.5 Lite model page lists three output rates: $0.02 per second at 480p, $0.03 at 720p and $0.14 at 1080p.
The full Video 1.5 model page lists $0.08, $0.14 and $0.25 per second at the same respective resolutions. The overview’s “from $0.08” price therefore should not be treated as the full model’s rate at every resolution.
| Resolution | Video 1.5 per second | Video 1.5 Lite per second | Video 1.5, 10-second output | Lite, 10-second output |
|---|---|---|---|---|
| 480p | $0.08 | $0.02 | $0.80 | $0.20 |
| 720p | $0.14 | $0.03 | $1.40 | $0.30 |
| 1080p | $0.25 | $0.14 | $2.50 | $1.40 |
The example totals are calculated from the published output rates and exclude input charges.
Lite is less expensive at every matching resolution. Its strongest price advantage is at the lower resolutions: a 10-second 720p clip costs $0.30 in output charges, compared with $1.40 on the full model. At 1080p, the gap narrows to $1.40 versus $2.50.
There’s also a useful crossover: Lite at 1080p costs the same per second as the full model at 720p. Developers working within a fixed budget therefore have a choice between a larger upscaled deliverable and the fuller model’s generation controls at a lower output resolution.
Both model pages list a $0.01 image-input charge. The full model also explicitly says preset audio input is free. Applications using visual references need to account for those inputs rather than estimating the entire job from duration alone.
Native 1080p and Upscaled 1080p Are Different Deliverables
xAI’s video generation guide distinguishes the two models’ 1080p paths. The full Grok Imagine Video 1.5 model renders 1080p natively, while Lite reaches 1080p by upscaling 720p output. The overview specifically advertises native 1080p generation from text or an image on the full model.
A 1080p output label consequently does not describe an equivalent generation process across both tiers. Upscaling increases the output resolution, but it should not be treated as proof that the underlying scene was generated with the same detail as a native 1080p render.
That makes the full model the more relevant candidate to evaluate when native resolution is a delivery requirement. It does not prove that the full model produces better motion, more accurate subjects or fewer artifacts in every scene. Those are separate quality questions.
For preview clips or workflows that eventually publish below 1080p, Lite’s lower-resolution rates may be more useful than its upscaled option. Paying $0.14 per second for Lite’s 1080p output needs a reason beyond selecting the largest available setting.
The guide also reveals a shared implementation detail: text-to-video on both 1.5 variants first generates an image from the prompt, then animates it. Developers still submit one text-to-video request, and the intermediate image is not returned. This describes the documented workflow, not a disclosed architecture or an independent explanation of output quality.
Reference Controls Are the Full Model’s Bigger Differentiator
Resolution is only part of the upgrade. According to the Imagine overview, the full model supports up to 14 reference images and three voice references, plus pinned first, last and mid-video frames.
Those controls address a different problem from inexpensive prompt-only generation: specifying what needs to remain consistent and where particular visual states should occur. A production workflow might use references to guide a subject’s appearance, while pinned frames provide constraints for the beginning, an intermediate moment or the ending.
These are documented controls, not guarantees that every generated frame will follow the references correctly. Their value depends on how reliably the model respects them in the developer’s actual material.
The supplied documentation positions Lite around text-to-video and image-to-video generation. It does not establish feature parity with the full model’s multi-reference and pinned-frame workflow. Developers should not assume that swapping the model identifier preserves every advanced request option.
Audio needs similar care. xAI says generated videos include audio by default, and its overview advertises lip-synced speech on both variants. On the full model, reference-to-video supports up to three preset speaker voices.
The documented voice inputs are preset voices. That should not be read as evidence that the API accepts arbitrary recordings for voice cloning. Default audio also means a generated asset should not automatically be treated as silent footage when planning an editing or publishing pipeline.
Duration Limits Depend on the Workflow
The generation guide documents a configurable duration of 1 to 15 seconds, with an eight-second default, for general video generation. The overview also explicitly lists up to 15 seconds for Lite.
Applications should set duration deliberately. Leaving it unspecified means budgeting around the documented default, rather than assuming the API chooses the shortest or cheapest output.
The same documentation contains limits for older workflows that should not be confused with the new model names. Reference-to-video on grok-imagine-video is capped at 10 seconds. Video editing retains the source video’s duration, with a documented maximum of 8.7 seconds, rather than accepting a custom duration.
Those statements concern the original grok-imagine-video model. They are not evidence of a blanket 10-second cap on Video 1.5 reference generation.
The guide also directs video editing and extension to that original model. For existing integrations, this is a reason to keep capability checks tied to both the model identifier and the operation. The two new generation models should not be treated as automatic replacements for every Imagine video endpoint.
Submit a Job, Then Poll for the Result
Video generation remains asynchronous. A REST request starts the job and returns a request_id; it does not immediately return a finished video.
The documented workflow is:
- Submit a generation request to
POST https://api.x.ai/v1/videos/generations. - Save the returned
request_id. - Poll
GET https://api.x.ai/v1/videos/{request_id}until the job completes or fails.
Sources
- Imagine API overviewdocs.x.ai
- Grok Imagine Video 1.5 Lite model pagedocs.x.ai
- full Video 1.5 model pagedocs.x.ai
- video generation guidedocs.x.ai





