Skip to main content
AI-generated conceptual illustration separating a ribbon-over-sea image into subject, motion, framing and lighting layers.
2026/08/09

Video to Prompt: Build a Clear AI Video Prompt from a Reference

Turn reference footage into a usable AI video prompt with an observation sheet, worked text-to-video and image-to-video examples, and a troubleshooting table.

A useful video-to-prompt workflow starts with what happens in the reference, then decides what to preserve in a new generation. “Cinematic, emotional, high quality” does not tell a model what the subject should do.

This guide uses an original exercise: a red ribbon tied to a pier rail moves in the sea breeze. The examples below are prompt drafts, not tested model outputs.

1. Write a small observation sheet

QuestionOur intended shot
Main subject?One red ribbon tied to a weathered wooden rail.
Action?The loose end lifts in the breeze and settles.
Setting?A pier with the sea behind it.
Camera?A fixed close view of the knot and loose end.
Light?Soft light on an overcast morning.
What must stay stable?The knot stays attached to the same rail.

If you are observing an actual video, keep uncertain details out of the factual column. You do not need to guess which camera brand filmed it.

2. Start with one action

A vague first draft might be:

A beautiful emotional red ribbon by the ocean, epic cinematic movement, amazing detail.

A clearer text-to-video draft is:

Close view of a red ribbon tied around a weathered wooden pier rail. The loose end lifts in the sea breeze, twists once, and settles against the wood. The knot stays attached. Soft overcast morning light, sea in the background, fixed camera.

This version names the subject, movement and framing. It does not ask for an establishing shot, a close-up, a storm and a sunset in the same short clip.

3. Adapt it when you already have a reference image

If the starting image already establishes the ribbon, rail, light and composition, try a shorter image-to-video draft:

The ribbon’s loose end lifts in a gentle breeze, twists once, and settles against the rail. The knot stays attached. The camera remains fixed.

Runway’s Gen-4 prompting guide recommends simple motion descriptions, positive phrasing and iteration for that model. This is model-specific guidance, not a universal rule for every video tool.

Check the documentation for the model you actually use. Duration, reference images and aspect ratio may belong in interface controls rather than in the text prompt.

4. Keep the editing plan separate from one generation

Suppose your reference sequence has a wide pier view, a ribbon close-up and a person taking it away. Save that as a three-shot plan. Generate or film each shot in a way your chosen tool supports, then edit them together.

A detailed reconstruction can record time ranges and sound notes for your planning. Before pasting it into a generator, simplify the portion intended for one clip. A timeline is useful to an editor; it does not mean a model will obey second-by-second instructions.

5. Change one thing when a result goes wrong

What you seeFirst revision to try
Camera moves, but you wanted a still shot.Remove competing camera instructions; state a fixed camera.
Ribbon detaches from the rail.Reduce the motion and restate that the knot stays attached.
Almost nothing moves.Name the loose end and its action clearly; check the input image permits that movement.
The shot changes setting halfway through.Remove scene transitions and keep one setting per test.
Motion works, but the look is wrong.Keep the motion wording and adjust one lighting or style detail.

These are troubleshooting experiments, not guaranteed fixes. Save the prompt, model settings and result together so you can tell what changed.

Use VidBreak to prepare the written draft

VidBreak’s video-to-prompt workflow reconstructs subjects, actions, timing, camera and sound from an authorized video range of up to 10 minutes. It creates written AI video prompts; rendering happens in the tool you choose.

Start by checking the reconstructed action against the video, then adapt the prompt for your model. The tool cannot recover a hidden original prompt with certainty. See the three-format sample or the prompt-result guide.

FAQ

Should I include sound in every video prompt? Only when your chosen model supports the sound you need. Otherwise keep sound as an editing note or produce it separately.

Is a longer prompt more accurate? Not automatically. Add details that remove a real ambiguity. Conflicting movements and multiple scene changes can make the intended shot less clear.

Can I paste an entire ten-minute reconstruction into a video generator? Treat it as a planning document. Break it into supported shots or clips and adapt each prompt to the destination model’s limits.