How to Write Text-to-Video Prompts That Work
A practical text-to-video prompt guide: the five things every prompt needs, what to delete, and how to draft cheaply before you render at 1080p.

Most bad AI clips are not a model problem. They are a text-to-video prompt problem — a pile of adjectives with no subject, no camera, and no light. The models are good enough now that the prompt is the variable you actually control.
Here is the way we write prompts on Video Gen Now, and the reason each part is there.
Why keyword soup stops working
Image models trained us to write like a tag list: cinematic, 8k, ultra detailed, masterpiece, trending. Video models want something different, because they have to resolve time as well as pixels. A tag list tells the model what the frame should look like, but nothing about what happens between frame one and frame 120 — so the model invents the motion, and its invention is usually a slow zoom on nothing.
Write the shot the way you would describe it to a camera operator who cannot see your reference. Two or three sentences beat twenty comma-separated words.
The five parts of a shot
Every prompt that survives contact with a render has these five things. Miss one and the model fills the gap for you.
- Subject — one clear thing the shot is about. "A faceted glass bottle", not "a product".
- Action — what changes during the clip. Steam rises, a hand enters frame, rain starts. If nothing changes, you get a still image that costs video credits.
- Camera — the move and the framing: slow push in, handheld follow, locked-off wide, overhead. This is the single highest-value word group in the prompt.
- Light — direction and quality: one hard light from the left, soft window light, neon spill on wet asphalt. Light is what makes a render look shot rather than generated.
- Look — the grade and format: 35mm, shallow depth of field, muted teal, high contrast black and white.
Before
glass bottle, product shot, cinematic, 8k, beautiful lighting, professional, ultra realistic
After
A faceted glass bottle stands on wet stone. The camera pushes in slowly as one hard light sweeps across the label, throwing a long shadow to the right. Deep blacks, shallow depth of field, 35mm.
Same subject. The second one tells the model when to move, where the light comes from, and what the frame should feel like — so there is nothing left for it to guess.
What to delete
- Resolution and quality words.
4K,8K,masterpiece,best qualitydo nothing — you pick the real resolution in the render settings. They just dilute the tokens that matter. - Negatives. "No people, no text, no blur" often summons the thing you excluded. Describe the frame you want instead.
- Two shots in one prompt. "She walks in, then we cut to the street" gives you a confused morph. One prompt, one shot. Cut them together afterwards.
- Stacked styles. "Pixar meets film noir meets watercolour" averages out into mush. Commit to one look.
Prompting changes with the mode
The three modes on the generator want genuinely different prompts, and this is where most credits get wasted.
Text to video — the full five-part shot above. You are describing a frame that does not exist yet, so nothing is implied.
Image to video — describe only the motion. Your image is already the first frame; re-describing the scene fights the picture you uploaded and often makes the model redraw your product. "The camera pushes in slowly as steam rises from the cup and the light shifts across the table" is a complete image-to-video prompt.
Reference to video — upload the subjects that have to stay recognisable, then name them positionally: "Image 1 is the woman, Image 2 is her dog. They walk together through a sunlit garden, handheld camera following behind." Calling them out by number is what keeps a face or a package consistent across the clip.
Draft cheap, ship expensive
Prompting is iteration, and iteration is where budgets die. Credits track real GPU cost, which scales with pixels — so the same prompt costs very different amounts depending on where you render it.
For a five-second 16:9 clip on Seedance 2.0:
- 480p — about 68 credits
- 720p — about 152 credits
- 1080p — about 341 credits
- 4K — about 778 credits
That is a 5× spread between your draft and your delivery. Write the prompt, run three or four variations at 480p until the motion and the framing are right, then re-render the winner once at 1080p. You will spend roughly a third of what you would have spent iterating at full resolution.
Two related things worth knowing:
- Aspect ratio moves the price too. At the same resolution, 21:9 is more pixels than 1:1, so it costs more — a five-second 1080p clip runs about 447 credits at 21:9 and about 192 at 1:1. Vertical 9:16 and wide 16:9 cost the same, so rendering both cuts of a campaign is symmetrical.
- Failed renders are refunded. If a job errors, the credits go straight back to your balance and the task stays in your library with the error, so you can adjust the prompt rather than re-typing it.
Questions people ask
How long should a text-to-video prompt be?
Two to four sentences. Long enough to cover subject, action, camera, light and look; short enough that no single instruction gets buried. Past roughly 80 words the later clauses start losing influence.
Should I write prompts in English?
Yes, where you can. The models have seen far more English training data, so English prompts give them the most to work with — even when the clip itself has no dialogue.
Why does my clip barely move?
You almost certainly wrote a description of a photograph. Add an explicit action and an explicit camera move. "Locked-off wide, dust drifting through a shaft of window light" moves; "beautiful empty room, cinematic" does not.
Can I reuse a prompt that worked?
Every render is stored next to the prompt that produced it, so you can reuse or fork a shot from your library instead of retyping it. Change one clause at a time — that is the only way to learn what a word is actually doing.
Takeaways
- Describe a shot, not a mood board: subject, action, camera, light, look.
- Match the prompt to the mode — motion only for image to video, numbered subjects for reference to video.
- Iterate at 480p, deliver at 1080p, and let the refund policy absorb the failures.
Ready to try it? Open the text-to-video generator, paste one of the five-part prompts above, and swap in your own subject. First render, first clip.