Sora Prompts for Product Video: A Framework That Actually Works (10 Prompt Templates)
Last December, a Shopify store selling ceramic incense holders ran two ads for the same product. One was a studio photo they paid $400 for. The other was a seven-second Sora clip of the holder on a wood table, smoke curling up through window light, prompted in under ten minutes. The Sora ad ran at 2.3x the click-through for three weeks before the algorithm cooled on it. The thing that sold the clip was not the lighting. It was that the smoke behaved like smoke.
Toys "R" Us ran a Sora brand film in 2024 and Shopify demoed partner workflows throughout 2025, but for small sellers the thing that mattered was never the big launches — it was learning how to write a Sora prompt that produced a shot you could use. This is the framework I settled on after roughly 400 Sora clips across cosmetics, fashion, food, and jewelry. Five elements, ten templates, a workflow. One honest note: as of October 2026, OpenAI retired the Sora consumer product on April 26, 2026 and shut down the Sora 2 API on September 24, 2026 (see the Sora 2 launch post and the deprecation notice). The framework still works — it transfers one-to-one to Veo 3, Kling, and Runway Gen-3, which were trained on prompts in the Sora grammar.
Why Sora works for product video (and when it doesn't)
The research paper OpenAI published alongside the original Sora, Video generation models as world simulators, buried the useful part in section three: Sora was trained to predict future frames conditioned on a physics-consistent world model. Translation: prompt "pour milk into a glass" and the milk behaves like milk — viscosity, splash, foam. Most 2025 video models gave you milk that fell through the glass.
That is why Sora dominated product briefs for eighteen months. Steam from food, condensation on cold cans, fabric drape, ice melting on a watch face, powder foundation puffing when you tap the compact — the shots that sell on Instagram, the shots other models faked badly.
Failure modes were just as consistent. Sora could not render legible text; packaging letters came out as nonsense. Close-up fingers-on-small-object shots had a 30% retry rate even in Sora 2. Hair in motion wobbled. Sora 2 maxed at 25 seconds with usable continuity dropping after about 15 — beyond ten, backgrounds drift. Practical rule: 3-to-8-second hero shots where physics matters, cut away before continuity breaks, no brand logos in frame unless you planned to composite them in post.
The five-element Sora prompt framework
Every prompt that produced a usable shot had the same five parts, roughly in order. Miss one and the model guesses, usually worse than you would.
1. Subject. What is in frame. Specific about material, color, size, condition. "A ceramic incense holder" is weak. "A matte-black ceramic incense holder, fist-sized, with a thin wisp of smoke rising from its center" resolves.
2. Motion. What moves and how. Writers describe scenes like photos. Sora needs a verb. "The smoke rises, curls to the left, drifts out of frame." Skip this and Sora invents motion, usually unwanted camera shake.
3. Camera. Lens, angle, movement. Sora was trained on film sets and reads camera-direction words well. "Slow dolly-in from eye level, 35mm lens, shallow depth of field." Skip it and you get a flat mid-shot that looks like a listing photo in motion.
4. Lighting. Source, direction, quality. "Soft window light from the left, late afternoon, warm tungsten fill from the right." Cinematography terms — Rembrandt, rim, practical, bounce — all work. "Good lighting" does not.
5. Style. Film stock, era, mood, grade. "Shot on 16mm Kodak Vision3 500T, slight grain, muted autumn palette" produces a wildly different result from "clean 4K digital, high key, pastel pink." Both valid. Blank is not.
Write these five lines as separate sentences, not a paragraph. Sora parses the structure and gives each element roughly equal weight; in a dense paragraph the first clause gets most of the attention and the rest is window dressing. The prompt optimizer on PromptSpace enforces this structure by default.
10 copy-paste Sora prompts for common product shots
Ten shot types that come up in e-commerce briefs. Each has a prompt, expected output, and one knob to turn when it fails. Replace bracketed sections with your product.
1. Cosmetic flat-lay with slow dolly
Prompt: "A [matte terracotta lipstick tube] lies on a cream linen surface alongside dried eucalyptus and a brass mirror. The lipstick stays still. Slow dolly-in from above, 50mm lens, shallow depth. Soft morning window light from the upper left. Shot on 35mm film, slight grain, muted earthy palette."
Expect: Six seconds of camera creeping toward the product. Fix: If the camera wobbles, add "camera on a motorized slider, no handheld shake."
2. Fashion model walk, 3-second loop
Prompt: "A woman in a [long oatmeal wool coat] walks away from the camera down a wet cobblestone street at dusk. Hair in a low bun; face not visible. Steady tracking shot behind her, 85mm lens, shallow depth. Streetlights come on one by one. Digital, cinematic grade, teal shadows and amber highlights."
Expect: Three to five usable seconds. Fix: If the walk glides, add "natural footstep rhythm, slight shoulder sway." Hair wobble is unfixable — cut before second four.
3. Food close-up with steam reveal
Prompt: "A ceramic bowl of [spicy ramen with soft-boiled egg and scallions] sits on a dark wood table. Steam rises in thick curls catching backlight. The bowl stays still. Slow push-in, 35mm macro lens, 24fps. Hard rim light from behind, deep shadows in front. Shot on 16mm Kodak Vision3 500T, warm highlights, dense blacks."
Expect: The shot Sora was born for; steam physics carries the clip. Fix: If steam looks fake, add "visible air currents, steam diffuses as it rises." If the yolk looks solid, specify "glossy, slightly running yolk."
4. Electronics unbox, hand reach
Prompt: "A matte-black cardboard box sits closed on a concrete surface. A hand enters from the right, lifts the lid slowly, and reveals a [brushed-aluminum wireless earbud case] inside. Overhead shot, 50mm lens, locked off. Single soft overhead key, slight falloff to the edges. Digital, clean modern grade, cool neutral tones."
Expect: Six seconds. Hands are the risky element — expect two or three retries. Fix: If fingers mesh or warp, re-run with "hand enters only partially from the right edge, fingers not fully visible."
5. Skincare pour onto skin
Prompt: "A single drop of [amber facial oil] falls from a glass dropper onto the back of a hand. The oil beads up on the skin, then slowly spreads. Extreme macro, 100mm lens, shallow depth. Soft diffused side light, high key, clean white background. Digital, clinical modern, crisp whites and warm skin tones."
Expect: Four to six seconds, droplet physics as the hero. Fix: If the drop looks like a CGI sphere, add "liquid viscosity, surface tension visible."
6. Jewelry 360 rotation
Prompt: "A [gold signet ring with a flat onyx stone] rotates slowly on matte black velvet. Camera locked off; the ring turns on an invisible turntable. 100mm macro, deep shadows. Three-point lighting: soft key upper left, hard rim from behind, subtle fill from below. Digital, luxury product grade, deep blacks and warm metallic highlights."
Expect: Clean 360 in about seven seconds; reflections shift realistically. Fix: If the ring judders, specify "constant angular velocity, 45 degrees per second."
7. Furniture room reveal
Prompt: "A [bouclé armchair in cream white] sits in a sunlit living room with a large window on the left and a woven rug underneath. Slow dolly-in from the doorway, 24mm wide, deep focus. Late afternoon golden-hour light through the window, dust motes visible. 35mm film, warm nostalgic grade, slight grain."
Expect: Eight seconds. Geometry holds for six before the far wall drifts — cut there. Fix: If the room looks like a stock image, add "lived-in details: a book on the floor, a half-full coffee cup on the side table."
8. Fitness equipment in use, face obscured
Prompt: "A pair of hands grips a [knurled steel kettlebell handle]. The kettlebell swings up into frame and out again. Low angle, camera at floor level, 35mm lens. Harsh overhead gym light, sweat visible on forearms. Face and head not in frame. Digital, high-contrast grade, cool shadows and warm highlights."
Expect: Five seconds, two full swings. Fix: If the kettlebell distorts mid-swing, slow it: "kettlebell swings slowly, half speed."
9. Packaged food, twist-off cap
Prompt: "A [glass bottle of cold-brew coffee] on a kitchen counter. A hand enters from the right and twists off the metal cap. Condensation beads on the glass. Eye-level, 50mm, shallow depth. Soft window light from the right, cool morning tones. Digital, naturalistic grade. No text visible on the label."
Expect: Four to six seconds; condensation is the sell. Fix: If the cap unscrews without a hand touching it, re-specify "hand grips cap, fingers wrap around the ridges, twists counter-clockwise."
10. Car exterior tracking shot
Prompt: "A [matte-charcoal electric SUV] drives slowly along an empty coastal highway at golden hour. Camera tracks alongside from a parallel vehicle, matched speed. Low angle, 35mm lens. Hard warm sunset from the right, long shadows across the road. Anamorphic lenses, cinematic grade, teal and orange."
Expect: Eight seconds of parallel tracking. Fix: If wheels appear static or warp, add "wheels spin at speed, motion blur on spokes."
Comparison: Sora vs. Runway vs. Kling vs. Veo for product video
Honest cost and capability snapshot from my notes between January and September 2026, before Sora wound down. Per 5-second 1080p clip. For a deeper head-to-head, we maintain a running text-to-video 5-second benchmark.
| Model | Cost per 5s | Hands | Camera direction | Duration cap | Best for |
|---|---|---|---|---|---|
| Sora 2 (deprecated) | ~$0.50 | Good | Excellent | 25s | Physics-heavy hero shots |
| Runway Gen-3 | ~$0.75 | Fair | Good | 10s | Stylized, moody looks |
| Kling 1.6 | ~$0.30 | Fair | Fair | 10s | Budget volume work |
| Veo 3 | ~$0.60 | Very good | Excellent | 8s | Realism with synced audio |
Veo 3 is the practical Sora replacement for most briefs. Kling is where you go for twenty variations. Runway is where you go when the brief says "cinematic" more than once.
Pro tip: Sora, Veo, and Runway all read camera-direction vocabulary from film sets. "Dolly," "pan," "tilt," "crane," "track," "handheld," "locked off" — these words produce the specific motion you asked for. Vague directions like "the camera moves around" produce whatever the model feels like, usually a slow unmotivated zoom. Write like a cinematographer, not a travel blogger.
Did you know: Sora's edge on physics was not an accident. The original Sora research paper describes training on data curated for consistent object permanence, gravity, and fluid behavior. That is why ice melting on a glass or water pouring into a cup looked right in Sora when every other 2024 model treated liquids as a texture. The successor models that work today — Veo 3 especially — inherited the same training approach.
How to iterate — the practical workflow
The biggest mistake was rewriting the entire prompt when a clip failed. This burns credits and teaches you nothing. The workflow that works:
1. Run the prompt once. Write down the one thing most wrong. 2. Change exactly one of the five elements — leave the other four alone. 3. Run again, compare. 4. If better, keep the change and move to the next issue. If worse, revert and try a different fix on the same element.
Three iterations of single-variable changes beat ten total rewrites. After six or seven generations you have a prompt you can reuse for the entire product line. Save the final version with your SKU in a text file. Build a library over the quarter.
Pricing and access reality in October 2026
Sora 2 is no longer accessible — consumer app retired April 26, 2026, API shut down September 24, 2026. Port your 2025 Sora prompts to Veo 3 or Kling; the grammar is nearly identical. Veo 3 runs ~$0.60 per 5-second 1080p clip through Google's video API. Kling 1.6 is the budget option at ~$0.30. Runway Gen-3 is $12-28/month with credit-based generation. A small Shopify store doing one launch a month can run Kling plus occasional Veo for under $40/month. Zero budget: Kling's free tier gives five generations a day with watermark.
Honest limitations
None of these models render legible text yet. If your product has a label, mask it in the prompt ("no text visible on packaging") and composite the real label in post. Hands are better than they were but still unreliable in close-up — shoot partial hands or none. Continuity past ten seconds breaks: backgrounds drift, props vanish, outfits change. For a 30-second ad, generate three 8-second clips and cut between them. And faces are a legal minefield; successor models do not permit named-person likeness without licensed training data. Keep faces obscured, blurred, or out of frame for anything commercial.
FAQ
Is Sora still available in October 2026?
No. OpenAI retired the consumer product on April 26, 2026 and shut down the Sora 2 API on September 24, 2026. The framework in this post still works because Veo 3, Kling 1.6, and Runway Gen-3 all respond to Sora-style prompts.
What is the best Sora alternative for product video today?
Veo 3 for physics-heavy shots and synced audio. Kling 1.6 for cheap volume work. Runway Gen-3 for stylized moods. Most product teams I know run Veo 3 as primary and Kling as the drafting tool.
Can I use these prompts commercially?
The prompts, yes. The outputs depend on the model's terms. Veo 3 and Kling allow commercial use on paid tiers. Check the plan before shipping an ad. Do not generate recognizable celebrity faces or copyrighted characters — that is where takedowns come from.
How long should a product video clip be?
For social, 3-8 seconds per shot. Cut between generated clips rather than asking one model for a 30-second continuous take.
Why does my Sora-style prompt produce different results in Veo 3?
Veo weighs audio and camera more heavily than Sora did. Add a specific audio direction ("ambient room tone, no music") and specify camera equipment ("Arri Alexa, 50mm Zeiss prime"), and the output tightens up fast.
Do I need image-to-video to show my actual product?
Yes, if you want the real product in frame. Text-to-video approximates it. Upload a clean reference photo and prompt the motion. All three successor models support image-to-video.
What is the one word that changes a prompt most?
The camera verb. "Dolly" versus "pan" versus "handheld" produces three different clips from the same subject. Pick the verb first, then write the rest. For more worked examples see the PromptSpace blog.
One more thing
The framework outlasts the model. Sora is gone, Veo is here, something else will be here in six months. Subject, motion, camera, lighting, style — that is how cinematographers have described shots since the 1930s. Write prompts the way a DP would describe a shot to a camera operator and you get what you asked for from whatever model is current. Go make something.












