Video Gen Now

Text to video

No footage, no camera, no set. Describe the shot the way you would describe it to a cinematographer — subject, action, camera, light — and the model shoots it. Seedance 2.0 adds a synchronized audio track; MiniMax H3 renders at 2K.

Describe the shot: subject, action, camera move, light. Two or three sentences beat one long list of keywords.

Native audio

Sound effects, ambience and lip-synced speech, generated from the same prompt. Costs the same either way.

152 credits

Rendering is paid in credits, and new accounts start at zero. Pricing

Made with this model

Every clip below was rendered here, on the settings shown under it. The caption is the exact prompt that produced it.

A ceramic pour-over kettle tips over a glass carafe and coffee falls in a thin dark ribbon; steam curls through a shaft of morning window light. Slow push-in, shallow depth of field, warm film grade.

Seedance 2.0 · 720p · 16:9 · 4s

A lone red kayak cuts across a glassy alpine lake at dawn; mist lifts off the water and the mountains catch the first light. Slow aerial orbit, cinematic wide shot.

MiniMax H3 · 2K · 16:9 · 5s

How text to video works here

  1. Write the shot

    Two or three sentences covering what is in frame, what moves, and how the camera behaves. Skip the keyword soup — these models read prose.

  2. Frame it

    Pick an aspect ratio for wherever the clip is going: 9:16 for vertical feeds, 16:9 for embeds and YouTube, 21:9 when you want it to feel like film.

  3. Render

    Four to fifteen seconds per clip. Progress streams live, and the finished MP4 downloads straight from the panel.

Specs at a glance

ModelDurationResolutionNative audioReference imagesFrom
Seedance 2.04–15 seconds480p · 720p · 1080p · 4KYesNot used38 credits for 5 seconds
MiniMax H35–15 seconds2KNoNot used130 credits for 5 seconds

What separates a good prompt from a lucky one

Lead with the subject, not the style. “A ceramicist trimming a bowl” anchors the model; “cinematic, 8k, masterpiece” does not.

Describe one camera move and commit to it. Competing moves in one prompt produce a clip that cannot decide.

Say what time of day it is. It sets colour, contrast and shadow direction in a single word.

If you want sound, say what you hear. Seedance 2.0 scores the clip from the same prompt — ask for rain on a tin roof and you get rain on a tin roof.

Text to video questions