Lumen AI logoLumen AI
AI Video Tools

How to Prompt AI Video Models: Shots, Motion and Continuity That Hold Up

A practical prompting guide for Sora, Runway, Veo and Kling — camera language, motion control, shot lists, and the continuity tricks that make generated clips cut together.

Lumen AI Editorial8 min readEdit this article
Filmmaker reviewing AI-generated video shots on a monitor with a storyboard beside the desk

Text-to-video models reward a completely different prompting style from image models. An image prompt describes a frame. A video prompt has to describe a frame and what changes over time — and if you leave the change unspecified, the model invents drift, warping, and a slow zoom nobody asked for.

Here is the structure that produces usable clips.

Storyboard sketches and shot list laid out next to a laptop running a video generator
A shot list beats a paragraph of description every time.

The five-part video prompt

Write prompts in this order. It maps to how a director briefs a crew, and models respond to it well.

  1. Shot type and subject — "Medium close-up of an older woman in a wool coat"
  2. Action, with a verb and a duration — "she slowly turns to look off-camera left"
  3. Camera behaviour — "static tripod shot" or "slow dolly in, 20mm"
  4. Environment and light — "rain-slick city street at night, sodium streetlights, shallow depth of field"
  5. Look — "shot on 35mm film, natural grain, muted colour"

One action per clip. The single most common mistake is stacking three beats into one generation: the model tries all of them, and the result is a smear.

Camera language the models understand

These terms produce consistently distinguishable results across the major models:

  • Static / locked-off shot — the most reliable way to avoid unwanted drift.
  • Slow dolly in / out, truck left / right, crane up, tilt down — deliberate, predictable moves.
  • Handheld, subtle shake — adds realism to documentary-style footage.
  • Wide, medium, close-up, extreme close-up — controls framing far better than "zoomed in".
  • Focal lengths — 24mm for environment, 50mm for natural, 85mm for portrait compression.

Avoid "cinematic" as a keyword. It is the "beautiful" of video prompting: everything and nothing.

Abstract representation of frames being generated in sequence by a model
Motion coherence is the hardest thing these models do.

Controlling motion

Three levers, in order of reliability:

Duration awareness. Most models generate 5–10 second clips. Ask for an action that genuinely fits that window. "Walks across the room and sits down" does not fit five seconds; "takes two steps toward the window" does.

Motion strength settings. Where the tool exposes a slider, low values plus a described camera move give cleaner results than high values with a vague prompt.

Image-to-video. Generate or shoot a still you already like, then animate it. This is the highest-control workflow available and it solves half of all composition problems before motion is involved.

Continuity across shots

This is where most projects fall apart. Adjectives do not carry a character across clips — references do.

  • Use a reference image or character feature where the tool supports it, and reuse the exact same file for every shot.
  • Freeze your wardrobe and lighting description as a copy-pasted block appended to every prompt. Change only the shot line.
  • Shoot around the problem. Over-the-shoulder, hands, objects, and environment cutaways carry a scene without demanding facial consistency.
  • Cut on motion. A match cut hides a lot of inconsistency that a slow dissolve exposes.

Professional teams treat this as coverage planning, not prompting: generate a wide, a medium, an insert, and a reaction for every beat, then edit.

Editor arranging generated clips into a sequence on a timeline
Generate more clips than you need and cut ruthlessly.

What still breaks in 2026

Be realistic about the failure list so you can plan around it:

  • Hands and fine manipulation — improving, still risky in close-up.
  • Legible text — signage and screens come out wrong. Composite text in post.
  • Physics with contact — pouring, catching, cutting, footfalls on stairs.
  • Long takes — coherence degrades noticeably past about ten seconds.
  • Exact repetition — a specific product or logo must be composited, not generated.
  • Synced dialogue — lip-sync tools exist but a talking-head performance is still the weak point.

A shot list beats a script

For a 60-second piece, plan roughly 12–18 shots. Write each as a single line in a spreadsheet with columns for shot type, action, prompt, generation count, and the chosen take. Then generate four to six variations per shot.

Expect a keeper rate around one in four. That ratio is the real production math — budget generations, not hours.

Post-production is not optional

Raw model output rarely looks finished. The pass that makes it feel professional:

  1. Colour grade everything to one LUT so clips from different generations match.
  2. Add grain — a consistent grain layer hides model artefacts remarkably well.
  3. Sound design carries the illusion. Room tone, footsteps, cloth movement, and a music bed do more for perceived quality than another generation attempt.
  4. Speed ramp to trim the mushy start or end of a clip.
  5. Stabilise or crop where drift is visible at the frame edge.
Production team reviewing character continuity across shots on a shared display
Continuity is managed with references, not with adjectives.

Rights, disclosure and safety

Check your model's commercial terms before client delivery — they vary by plan and by region. Do not generate identifiable real people without consent, and label synthetic media where the audience could reasonably be misled. Several platforms now attach provenance metadata automatically; leave it intact.

The bottom line

Prompting video well is mostly directing: one action per shot, an explicit camera, a fixed look, and coverage planned in advance. The people getting broadcast-usable results are not writing better sentences — they are running a small production process.

Compare the models themselves in our AI video generator roundup, see the editing side in the AI video editing workflow, and browse more in AI Video Tools.

#Prompting#Sora#Runway#Veo#Cinematography