Writing prompts
The structure that turns a vague description into a shot.
Prompt quality changes results more than any other setting. A model given "a cat in a city" has to invent the breed, the city, the time of day, the lens, and the motion — and it will invent something different every run. Your job is to remove the guesses that matter to you and leave the rest.
The five-part structure
Most good prompts describe the same five things, in roughly this order:
- Subject — who or what, described concretely. An elderly fisherman in a yellow oilskin coat.
- Action — what changes during the clip. He hauls a net over the gunwale, straining against the weight.
- Setting — where, and when. On the deck of a small trawler at dawn, open grey sea behind him.
- Camera — how it is seen. Handheld medium shot from the side, slowly drifting closer.
- Look — light, colour, texture, style. Overcast light, desaturated blues, 35 mm film grain.
Put together:
An elderly fisherman in a yellow oilskin coat hauls a net over the gunwale,
straining against the weight, on the deck of a small trawler at dawn with open
grey sea behind him. Handheld medium shot from the side, slowly drifting
closer. Overcast light, desaturated blues, 35mm film grain.You do not need all five every time. But whichever one you leave out is the one the model will surprise you with.
Describe motion, not a photograph
The most common mistake is writing a still-image prompt. "A woman standing in a neon-lit alley" describes a photograph — nothing in it has to change, so the model produces something close to a moving still.
Give the clip a beginning and an end:
| Instead of | Write |
|---|---|
| A woman in a neon alley | She steps out of the doorway into the rain and pulls up her hood |
| A busy market | The vendor tosses fruit into a bag while the crowd moves past behind him |
| A spaceship | The ship banks hard left, thrusters flaring, and drops below the cloud line |
Direct the camera
Camera language does real work. Useful terms:
- Shot size — wide shot, medium shot, close-up, extreme close-up
- Movement — static, handheld, slow push in, pull back, tracking shot, crane up, orbit
- Angle — eye level, low angle, high angle, overhead
- Lens feel — wide angle, telephoto, shallow depth of field
One camera instruction per clip. Asking for a push in and an orbit and a crane in five seconds gives you an incoherent smear.
Name the light
Light is what makes a shot read as cinematic rather than as a render. "Golden hour backlight", "harsh overhead fluorescents", "single candle, deep shadows", "overcast soft light" each produce a completely different image from the same subject.
Match the prompt to the length
The clip length you chose is your budget for action. Roughly:
- 4–5 seconds — one action. A turn, a step, a door opening.
- 8–10 seconds — one action with a camera move, or two small beats.
- 15–30 seconds — a sequence. Give it structure, or it will drift.
Asking for three scene changes in five seconds is the fastest way to get mush. When you genuinely need a sequence, consider generating shorter shots and chaining them — see Image to video.
Say what you do not want, positively
Negative phrasing is unreliable — "no text on screen" can put text on screen. Describe the positive state instead: "clean unmarked walls" rather than "no graffiti", "empty street" rather than "no people".
Iterate cheaply
Prompt iteration should not be expensive. Draft at 480p and 4 seconds, change one thing at a time, and only re-run at full resolution and length once the shot is right. Changing three things at once tells you nothing about which one helped.
Remember that every run is a fresh sample. If a prompt produces a good result once and a bad one the next time, the prompt is fine — the sample was unlucky. Regenerate before rewriting.
Steal from the gallery
The Inspiration gallery shows generations alongside the prompts that produced them. Opening one, changing the subject, and keeping its camera and lighting language is a much faster way to learn what this model responds to than writing from scratch.