Google DeepMind shipped Veo 3.1 in late 2026, and the headline change isn't resolution or length — it's that the model finally does what you tell it. Prompt adherence jumped hard. Camera moves land. Subject placement holds. Lighting descriptions don't get quietly rewritten by the sampler. If you spent 2025 fighting Veo 3 to keep a character in frame, this update is going to feel almost boring in a good way.
The catch: to get the payoff, you have to write prompts like a shot list, not a wish. This guide covers the structure that works, four cinematic examples you can steal, how Veo 3.1 stacks up against Sora 2, Runway Gen 4, and Kling 2.5, and the habits that separate a usable clip from a re-roll.
What actually changed in Veo 3.1
Veo 3.1 keeps the native 1080p output and synchronized audio from Veo 3, but rebuilds the text encoder and adds a stronger motion prior. In practice that means three things you'll feel immediately. First, camera language is respected — say "slow dolly in" and you get a dolly, not a zoom. Second, negative prompts finally work; if you write "no lens flare," the model listens. Third, multi-subject scenes hold identity across the full 8-second clip instead of morphing at second five.
Length went up too. Veo 3.1 renders up to 12 seconds in a single pass at 1080p, or stitches to 60 seconds via the extend endpoint without color drift. Audio covers dialogue, ambient, and Foley in one prompt.
Pro tip: Write the shot in this order — subject, action, environment, camera, lens, lighting, mood, audio. Veo 3.1's encoder weights early tokens more, so lead with what has to be right.
The prompt structure that works
Every reliable veo 3.1 prompt I've shipped follows the same seven-part skeleton. It's not a rule, it's a scaffold — but skipping parts is where re-rolls come from. The parts are: subject (who or what, one sentence, specific), action (verb-first, present tense), environment (place, time of day, weather), camera (shot size and movement), lens and format (focal length, aspect ratio, film stock if you want a look), lighting and mood (source, direction, color temperature), and audio (dialogue, ambient, or "no dialogue").
Keep the whole thing under 90 words. Longer prompts don't help — the encoder truncates hard around 120 tokens, and the last thing you want is your camera direction getting dropped because you over-described the wallpaper.
Four cinematic prompt examples
These are real prompts, tested on Veo 3.1 in September 2026. Copy them, swap the nouns, and you'll see the pattern.
1. Neo-noir chase, rooftop
A woman in a wet trench coat sprints across a rain-soaked rooftop at 3 a.m., breath visible. Neon signs bleed pink and cyan across puddles. Handheld medium shot, 35mm anamorphic lens, shallow depth of field. Practical light from a flickering vent above. Cold, high-contrast, Blade Runner mood. Audio: hard footfalls on wet gravel, distant siren, no dialogue. 16:9, 10 seconds.
2. Wildlife documentary, golden hour
A red fox pauses at the edge of a snowy pine forest, ears forward, then trots left across frame. Low-angle static wide shot on a 24mm lens. Late-afternoon backlight through the trees casts long blue shadows on the snow. Warm 5600K key, natural. Calm, patient BBC nature-doc mood. Audio: crunching snow, distant crow, wind through pines. 16:9, 8 seconds.
3. Product hero shot, kitchen
A matte black espresso machine sits on a walnut counter as steam curls from the portafilter. Slow 6-second dolly in from wide to close-up on the drip. 50mm lens, f/2.8, cinematic. Warm morning window light from camera left, soft shadow on the wall. Editorial, minimal, premium. Audio: soft espresso hiss, ceramic cup placed down, no music. 9:16 vertical, 8 seconds.
4. Sci-fi establishing, desert
A lone figure in a dust-caked orange spacesuit walks toward a colossal derelict starship half-buried in red desert sand. Slow aerial pull-back reveals scale. 35mm lens, wide, deep focus. Twin-sun harsh top light, long shadows, hazy horizon. Awe, isolation, Denis Villeneuve tone. Audio: wind, distant metallic groan, muffled breath in helmet. 2.39:1, 12 seconds.
Did you know? Veo 3.1 was trained with a dedicated cinematography head that saw millions of annotated shot descriptions. That's why terms like "anamorphic," "dolly in," and "f/2.8" behave like actual controls, not vibes.
Veo 3.1 vs Sora 2, Runway Gen 4, Kling 2.5
Every model has a personality. Here's how the late-2026 lineup actually behaves when you push it.
| Model | Max length | Native audio | Prompt adherence | Best for |
|---|---|---|---|---|
| Veo 3.1 | 12s single / 60s extend | Yes, dialogue + Foley | Excellent | Cinematic, controlled shots |
| Sora 2 | 20s | Yes | Very good | Surreal, imaginative scenes |
| Runway Gen 4 | 10s | No (separate call) | Good | Motion graphics, VFX plates |
| Kling 2.5 | 10s | Limited | Good | Human motion, dance, sports |
Pick Veo 3.1 when the shot needs to match a storyboard. Pick Sora 2 when you want the model to surprise you. Runway is still the workhorse for anyone integrating video into an existing edit pipeline. Kling remains the human-motion king if you're doing anything with dancers or athletes.
Common mistakes and how to fix them
Three failure modes show up over and over in the Veo 3.1 subreddit and our own testing. Wall-of-text prompts — if it reads like a novel, the encoder is truncating your camera direction. Cut it to 80 words. Conflicting camera instructions — "static wide shot with a slow zoom" contradicts itself; the model will pick one and you won't like which. Choose static or moving, not both. Vague lighting — "cinematic lighting" means nothing. Say the source (window, practical, sun), direction (left, backlight), and temperature (warm, 5600K).
For dialogue, write it in quotes and specify the speaker: The old man says, "You're late," in a gravelly voice. Veo 3.1 will lip-sync it. Skip the quotes and you'll get muttered gibberish that sort of matches the mouth.
Frequently asked questions
Is Veo 3.1 free to use?
Google offers a limited free tier through the Gemini app and AI Studio; heavy use routes through Vertex AI with per-second pricing. You can also try prompts inside our free AI video generator without a Google account.
What's the maximum resolution?
1080p native. Veo 3.1 does not currently output 4K in a single pass, though the upscale endpoint gets you there for finishing.
Can Veo 3.1 do image-to-video?
Yes. Upload a still, add a motion prompt, and it animates from your frame while preserving identity better than Veo 3 did.
How long should a prompt be?
40 to 80 words is the sweet spot. Past 120 tokens the encoder starts dropping tail content.
Does it handle text in the scene?
Short text (signs, labels, license plates) works well. Long paragraphs of on-screen text still garble.
Can I control camera movement precisely?
Yes — dolly, truck, pan, tilt, crane, and orbit are all understood. Combine one movement per shot for reliability.
Where can I read the official spec?
Google DeepMind publishes the model card and capabilities at deepmind.google/technologies/veo.
What's the license for commercial use?
Outputs from Vertex AI are cleared for commercial use under Google's standard generative AI terms. Free-tier outputs have a visible watermark.
Start shipping shots
The gap between a mediocre Veo 3.1 clip and a great one isn't the model — it's how you write the prompt. Lead with the subject, name your camera move, pick one lighting source, and keep the whole thing tight. Do that and Veo 3.1 will render close to what you pictured on the first or second try, which is a genuinely new feeling in AI video.
Try our free AI video generator to test these prompt patterns in your browser, or browse more prompt guides on the PromptSpace blog for Sora 2, Runway Gen 4, and Kling breakdowns.












