LNKLING
Five recent video models
The video model is becoming a studio
The newest video models are no longer competing only to make a beautiful clip. From Seedance’s 30-second one-take, to Veo’s native audio, to Gemini Omni’s multi-turn conversational edits, video generation is being framed as a full way of working.
This brief compares five models by positioning, duration, sound, and working mode — so you can decide whether you need direction, composition, or editing first.
00Essay
From clip generation to production system
Calling these systems “text-to-video models” now misses what matters. Seedance 2.5 stretches a single generation to 30 seconds with further extension; Veo 3.1 meets video with audio; Gemini Omni turns creation and revision into conversation; MiniMax H3 tries to unify multimodal tasks in one context; Kling Video 3.0 aims at 15-second commercial continuity with native audio and multi-shot storytelling.
Their differences are increasingly differences in workflow. Before choosing, ask what kind of work you are doing: directing a continuous passage, generating a finished audiovisual short, revising existing media through language, or organizing image, video, and audio references.
Direct. A longer one-take window and reference control let action, camera, and sound be staged inside one passage.
Compose. Image, dialogue, ambience, and effects start to arrive as one audiovisual object.
Edit. Video can inherit context and be revised over several natural-language turns instead of full restarts.
—At a glance
At a glance
| Model | Positioning | Duration | Sound | Working mode |
|---|---|---|---|---|
| Seedance 2.5 | One-take direction | Up to 30 seconds per generation; extend twice | Joint audio-video generation | Multimodal reference, timestamp edits, white-model & green-screen |
| Veo 3.1 | Audiovisual realism | Supports extension; limits depend on product surface | Native audio (dialogue, ambience, music, effects) | Video generation in Gemini and Google Flow |
| Gemini Omni | Conversational editing | Gemini overview states about 10-second videos | Native audio generation | Multi-turn editing with text, image, and video |
| MiniMax H3 | Multimodal creation | Short-form generation; official page showcases 2K + stereo | Native stereo audio | Unified multimodal-context understanding and generation |
| Kling Video 3.0 | Commercial short-form production | Up to 15 seconds | Native audio with character-level lip sync | Multi-shot storytelling, element consistency, multilingual lip sync |
01Model
Seedance 2.5
ByteDance Seed · One-take direction
ByteDance Seed’s product page calls Seedance 2.5 a next-generation audio-video joint generation model built for 30-second storytelling, precise reference control, and powerful editing. The launch article frames the same release as one-take creation: moving from generating a clip to completing a creative work.
Duration. Up to 30 seconds per generation; extend twice
Sound. Joint audio-video generation
Working mode. Multimodal reference, timestamp edits, white-model & green-screen
The product page states that users can create videos up to 30 seconds in a single generation, with the option to extend twice for richer storytelling, alongside smoother motion and more realistic visuals. The launch article adds that single-pass generation grows from 15 to 30 seconds and supports multi-round extensions, so a story can unfold through setup, development, turning points, and resolution rather than stretching one moment.
Reference and editing form the second axis. Official material says the model understands reference video more precisely—intention, framing, and cinematic language—beyond simple motion transfer. The launch article lists up to 30 images, 10 video clips, and 10 audio clips in one pass, plus clay render, motion, and creative references. Editing includes timestamp-level audio-video control, green screen, camera perspective, and reference-based edits for film and advertising workflows.
For professional production, the product page highlights white-model control, green-screen editing, camera movement, and performance blocking. The launch article says Seedance 2.5 is rolling out on Jimeng AI, Doubao Pro, and other creative products. This brief records those public claims only; it is not an independent benchmark.
Key points: Up to 30 seconds per generation, with the option to extend twice (product page); Joint audio-video generation; multi-round extension for longer narratives (launch article); Up to 30 images, 10 videos, and 10 audio clips as references (launch article); Timestamp editing, white-model control, green-screen and camera edits (official pages).
Links: Product page, Launch article.
02Model
Veo 3.1
Google DeepMind · Audiovisual realism
DeepMind’s Veo page opens with a single line: “Video, meet audio.” Veo 3.1 is presented as Google’s leading video generation model for filmmakers and storytellers, with sound treated as part of the generated shot.
Duration. Supports extension; limits depend on product surface
Sound. Native audio (dialogue, ambience, music, effects)
Working mode. Video generation in Gemini and Google Flow
The official capability story focuses on greater realism and fidelity grounded in real-world physics and audio, stronger prompt adherence, and improved creative control that now spans audio. DeepMind’s own sample prompts describe shot scale, performance, city murmurs, a hip-hop bed, and spoken dialogue in one instruction—audio is authored with the picture, not attached later.
The DeepMind page points creators to Gemini and Google Flow. Gemini offers a direct conversational entry, while Flow is organized around shots and longer creative sequences, so the practical choice depends on the complexity of the project.
Relative to silent clip generators, Veo 3.1’s public difference is audiovisual unity: dialogue, ambience, and rhythm can arrive with the image. Exact duration caps, resolution options, and regional availability follow each surface’s current documentation; this article does not invent a single public limit where the pages do not state one.
Key points: DeepMind positioning: leading video model for filmmakers and storytellers; Emphasis on native audio, physical realism, prompt adherence, and creative control; Creative products include Gemini and Google Flow; Official demos write ambience, music, and dialogue into the same prompt.
Links: Google DeepMind.
03Model
Gemini Omni
Google DeepMind · Conversational editing
Gemini’s video overview and the Google Blog present Gemini Omni as video creation and editing that feels like a conversation. The claim is not only one-prompt generation, but multimodal input plus multi-turn revision in one workflow.
Duration. Gemini overview states about 10-second videos
Sound. Native audio generation
Working mode. Multi-turn editing with text, image, and video
The Gemini overview lists create ~10-second videos, native audio generation, turn photos into video (up to 5), video-to-video editing, and multi-turn editing. In the Gemini app, Omni is described as replacing the previous Veo 3.1 path. DeepMind’s model page adds the framing “create anything from any input—starting with video,” with natural language edits that each build on the previous state.
The Google Blog introduces Gemini Omni Flash as the first model in the family, rolling out to the Gemini app, Google Flow, and YouTube Shorts. Users can combine images, audio, video, and text as input and generate video grounded in Gemini’s real-world knowledge. Across turns, official copy stresses character consistency, physics that hold up, and a scene that remembers what came before.
The public differentiator is interaction: swap a background, change wardrobe, adjust lighting, or replace a character through chat instead of restarting from zero. This brief follows those official pages only and does not treat demos as controlled comparative tests.
Key points: Gemini overview: ~10-second videos, native audio, up to 5 photo references; Video-to-video editing and multi-turn conversational edits; Text, image, audio, and video can be combined as input; Rolling out to Gemini app, Google Flow, and YouTube Shorts (Google Blog).
Links: Gemini overview, Google DeepMind, Google blog.
04Model
MiniMax H3
MiniMax · Multimodal creation
MiniMax’s H3 research page presents a unified multimodal creative tool: describe relationships among text, images, video, and audio in natural language, then use those materials to generate video with native stereo sound, including 2K demos.
Duration. Short-form generation; official page showcases 2K + stereo
Sound. Native stereo audio
Working mode. Unified multimodal-context understanding and generation
The official example is concrete: borrow Hitchcock-style camera movement from a video, have a character from an image sing, and match the performance to a separate audio track. The creator describes how the materials relate instead of handling each reference in a different tool.
Earlier products often separated image animation, first-and-last-frame control, subject, motion and voice references, and editing into different workflows. H3 aims to keep them in one creative context, where multi-shot video, sound, references, and natural-language revisions can continue together.
The page also highlights 2K output and native stereo sound, with use cases spanning film titles, product showcases, animated posters, and advertising. For creators, the useful question is whether different kinds of source material can be understood and used together.
Key points: Unified text, image, video, and audio context (official research page); Showcases 2K output and native stereo sound; Relationships among references described in natural language; Suited to film titles, product showcases, animated posters, and other mixed-media work.
Links: MiniMax research.
05Model
Kling Video 3.0
Kuaishou · Commercial short-form production
Kling’s product page frames Video 3.0 as professional cinematic short-form production: native audio, multi-shot storytelling, character consistency, and precise lip-sync for multilingual marketing. The Omni user guide upgrades that into all-in-one multimodal input, voice-driven characters, direct audio-visual output, and storyboarding.
Duration. Up to 15 seconds
Sound. Native audio with character-level lip sync
Working mode. Multi-shot storytelling, element consistency, multilingual lip sync
The product page lists what creators can make: multi-shot cinematic shorts with dynamic camera movement, localized ads with characters speaking different languages and accents, and up to 15 seconds of high-fidelity continuous action. What’s new emphasizes cinematic audiovisual synergy, element and character consistency, and native text rendering inside 15-second sequences for ads, product demos, and commercial content.
The VIDEO 3.0 Omni guide contrasts O1 with the new model: text- and image-to-video gain native audio and multi-shot support; video element reference and element voice control are added; duration rises from up to 10s to up to 15s. Images, videos, elements, and text are all treated as prompts, and the model can lock multiple characters or items across complex group scenes.
Official marketing also presents cinema-grade native 4K as part of the series. For deliverable short-form work, the stated priorities are stable subjects, accurate lip movement, readable text, and continuous shots—not a single spectacular frame. Capability claims here come only from the Kling product page and Omni user guide.
Key points: Up to 15 seconds per generation; Omni guide notes the rise from 10s to 15s; Native audio with character-level lip sync; multi-person and multilingual marketing; Multi-shot storytelling, element/character consistency, video elements and voice control (Omni); Product page emphasizes native text rendering and cinema-grade native 4K.
Links: User guide, Product page.
—Choice
Choose the working mode first
You need a more complete continuous sequence with reference and timestamp control
Seedance 2.5Sound, speech, and image need to feel credible in one generation
Veo 3.1You begin with existing media and revise through conversation
Gemini OmniYou want to combine image, video, and audio references with 2K output and native stereo
MiniMax H3You make ads, multi-shot character content, or multilingual lip-sync marketing
Kling Video 3.0Capabilities, regions, pricing, and product surfaces change over time — check the vendor pages for the latest details.
—Further reading
Further reading
Continue with these vendor pages if you want more depth on a specific model.
- 01Seedance 2.5
- 02Veo 3.1
- 03Gemini Omni
- 04MiniMax H3
- 05Kling Video 3.0