A useful Wan 3.0 prompt is not the longest description you can write. It is a compact set of production decisions the model can follow from the opening frame to the final beat.
This guide gives you a repeatable way to make those decisions. It works for text-to-video, image-to-video, and reference-to-video, but the emphasis changes with the input mode. You will also learn how to revise a weak result without rewriting everything and losing the parts that already worked.
The short formula
For text-to-video, start with:
Subject + setting + action + camera + visual treatment + sound + ending beat
For image-to-video, the image already defines much of the subject, setting, composition, and style. Use a tighter formula:
Primary motion + secondary motion + camera behavior + preservation instruction + ending state
For reference-to-video, add an asset manifest:
Reference role + subject action + scene relationship + camera + sound + ending beat
These formulas are planning tools, not magic strings. Their value is that they expose missing decisions before you spend a generation.
Start with the shot’s job
Before describing lighting or lenses, finish this sentence:
“The viewer should understand or feel __ by the end of the clip.”
A product reveal might need the viewer to understand how a mechanism opens. A character moment might need the viewer to feel hesitation turning into resolve. A travel scene might need to establish scale and atmosphere.
If you cannot state the job in one sentence, the prompt probably contains several competing videos. Split the idea before adding more detail.
Build the director’s brief in eight passes
1. Name the subject precisely
Identify the main subject with two or three durable traits. Choose details that remain visible during motion: silhouette, clothing, material, color, age range, or product geometry.
Weak:
A woman walks through a market.
Stronger:
A middle-aged florist in a moss-green apron carries a shallow wooden tray of white peonies through a covered morning market.
The stronger version does not merely add adjectives. It gives the scene a recognizable subject, prop, palette, and environment.
2. Establish the scene
Describe the space in terms that affect the shot: scale, foreground, background, time of day, weather, and practical light sources. Avoid listing decorative objects that never matter.
Condensation hangs under the glass roof; narrow sunbeams catch mist above the flower stalls, while shoppers move softly in the distant background.
3. Choose one primary action
Write the action as a visible change. “Feels hopeful” is internal; “stops, notices the first sunbeam, and lifts her gaze” can be filmed.
For a single shot, one clear action is usually stronger than a chain of unrelated gestures. If you want several beats, allocate them on a timeline instead of joining them with a long series of “and then” phrases.
4. Add secondary motion
Secondary motion makes the frame feel alive without competing with the subject. Examples include fabric responding to wind, reflections moving across glass, steam rising, background traffic crossing slowly, or leaves shifting after the subject passes.
Keep the hierarchy explicit:
- Primary: the florist stops and raises her eyes.
- Secondary: apron ties and loose petals move in a light draft.
- Background: shoppers remain soft and unhurried.
5. Direct the camera
Choose one camera behavior that supports the emotional job:
| Camera choice | What it communicates | Useful wording |
|---|---|---|
| Fixed frame | Observation, restraint, product clarity | “Locked camera; no pan or zoom” |
| Slow push-in | Attention, intimacy, realization | “Slow, steady push from medium shot to close-up” |
| Pull-back | Reveal, scale, isolation | “Camera retreats gradually to expose the full space” |
| Lateral track | Movement through a world | “Track beside the subject at walking pace” |
| Gentle orbit | Form, status, product detail | “Small clockwise orbit while keeping the subject centered” |
Do not stack a crane, orbit, zoom, handheld shake, and whip pan into one short clip unless the shot is intentionally chaotic. Camera instructions compete just like character actions do.
6. Define the visual treatment
Use concrete image-making choices rather than a pile of prestige words. Specify only what changes the frame:
- soft overcast daylight or hard noon sun;
- restrained earth tones or saturated neon contrast;
- shallow focus or clear deep focus;
- observational documentary framing or polished product-film composition.
“Cinematic, beautiful, epic, award-winning” does not resolve a creative decision. “Low winter sun, cool shadows, warm practical lamps, gentle highlight roll-off” does.
7. Plan sound as part of the scene
Wan 3.0 supports audio-visual generation, so sound belongs in the brief rather than as an afterthought. Separate the layers:
- Voice: exact dialogue, language, delivery, pace, and speaker.
- Effects: footsteps, fabric, machinery, rain, room tone, or object impacts.
- Music: presence, absence, instrumentation, energy, and when it enters.
Example:
No dialogue. Soft market room tone, distant cart wheels, a brief rustle of paper around the flowers. No background music until the final two seconds, when one warm cello note enters.
If you do not want dialogue or music, say so directly. Silence is a direction.
8. State the ending beat
The ending should be a visible, holdable state—not simply “the video ends.”
End on a steady close-up as one loose petal lands on the tray; hold the composition for the final second.
This gives the generation somewhere to arrive and produces a cleaner edit point.
A complete text-to-video example
Single continuous shot. A middle-aged florist in a moss-green apron carries a shallow wooden tray of white peonies through a covered morning market. Condensation hangs under the glass roof, and narrow sunbeams catch mist above the stalls. She walks toward camera, notices the first beam of sunlight, slows, and lifts her gaze with a restrained smile. Track backward at her walking pace, then make a very slow push-in as she stops. Keep the background shoppers soft and unhurried; apron ties and loose petals move in a light draft. Natural cool daylight with warm highlights on the flowers, realistic materials, shallow depth of field. No dialogue. Soft market room tone and distant cart wheels; one warm cello note enters in the final two seconds. End on a steady close-up as one petal lands on the tray, and hold for one second.
Notice the order: subject, space, action, camera, secondary motion, visual treatment, sound, ending. A reviewer can point to any clause and decide whether it helped.
Use a timeline for longer clips
Longer duration gives an idea room to develop, but it also increases the cost of vague pacing. For a 15- or 30-second piece, divide the prompt into readable beats.
Example structure:
- 0–5 seconds — Establish: introduce the subject, location, and camera relationship.
- 5–12 seconds — Develop: perform the central action or reveal information.
- 12–18 seconds — Turn: change emotion, direction, scale, or sound.
- 18–24 seconds — Resolve: complete the action and simplify the frame.
- 24–30 seconds — Hold: land on an image that can cut cleanly.
Do not force five beats into every 30-second video. A continuous performance may need only an opening, development, and final hold. The timeline should clarify the idea, not make it busier.
Change the emphasis for image-to-video
In image-to-video, repeating every visible detail can pull the generation away from the source. Treat the image as the art department’s approved frame and direct what changes.
The model remains seated and keeps the same face, hairstyle, jacket, and framing. She takes one slow breath, turns her eyes toward the window, then slightly relaxes her shoulders. A narrow curtain edge moves in the breeze; the rest of the room remains still. Locked camera with no zoom. Preserve the original color palette and facial proportions. End with her gaze held toward the window.
The prompt focuses on motion, restraint, and preservation. For a deeper workflow, read the Wan 3.0 image-to-video guide.
Assign references by job
When using several assets, name them consistently and state what each one controls:
Image 1: lead character identity and wardrobe;Image 2: location layout and color palette;Video 1: walking pace and shoulder movement;Audio 1: rhythm reference only.
Then write relationships, not a loose inventory:
The character from Image 1 crosses the corridor from Image 2 with the measured walking rhythm from Video 1. Keep Image 1’s charcoal coat and Image 2’s green tile pattern. Use Audio 1 only to guide the cut rhythm; do not reproduce its voice.
More assets are not automatically more control. Remove any reference whose job you cannot explain in one line.
Revise one variable at a time
When a result is close, preserve the prompt and change one layer:
| Problem | First revision | What to keep fixed |
|---|---|---|
| Motion feels weak | Clarify speed, path, and amplitude | Subject, scene, camera, style |
| Face or product drifts | Reduce pose change and camera rotation | Lighting, duration, overall action |
| Frame feels chaotic | Remove background actions | Main subject and ending beat |
| Camera fights the action | Replace multiple moves with one | Subject motion and scene |
| Ending cuts abruptly | Add a defined final state and hold | Earlier beats |
| Sound feels crowded | Assign voice, effects, and music separately | Visual direction |
Rewriting the entire prompt after every result makes diagnosis impossible. A controlled revision loop teaches you which instruction changed the outcome.
Preflight checklist
Before generating, confirm that:
- the shot has one clear job;
- the subject can be identified quickly;
- the primary action is visible and physically readable;
- camera direction supports rather than competes with that action;
- sound instructions name voice, effects, and music separately;
- references each have one explicit role;
- the ending is a defined visual state;
- the prompt does not contain contradictory speeds, moods, locations, or camera moves;
- the selected duration and aspect ratio match the delivery channel.
Then open the Wan 3.0 video generator and save the prompt with the output. Good prompting becomes repeatable only when the brief, settings, references, and result stay together.


