Wan 3.0 prompt guide

Wan 3.0 Prompt Guide: Write Better Video Prompts

Turn an idea into a directable video brief. This practical Wan 3.0 prompt guide explains what to specify, what to leave out, and how to revise one decision at a time.

12-minute guideUpdated September 12, 202610 copy-ready prompts

01 · Anatomy

Treat the prompt as a director’s brief

This Wan 3.0 prompt guide starts with a simple principle: a strong prompt is not a bag of cinematic adjectives. It is a compact sequence of production choices. The model needs to know who or what holds attention, where the action happens, what visibly changes, how the camera observes that change, what the scene sounds like, and where the shot should finish. When those decisions are explicit, a reviewer can also identify which instruction failed and revise it without discarding everything else.

Begin with the job of the clip. An advertisement may need to prove how a mechanism opens. A UGC scene may need to make a reaction feel spontaneous. A short narrative may need to turn curiosity into relief. Write that communication goal privately before the prompt, then choose one primary action that makes the goal legible. If the prompt asks for several unrelated actions, locations, and emotional turns inside a short duration, split the concept into separate generations or organize it as a timed multi-shot sequence.

Use concrete language. “Premium” is an opinion; brushed aluminum, controlled edge highlights, a centered product, and quiet room tone are visible or audible choices. “Dynamic camera” is vague; a low lateral track that keeps pace with a runner is directable. Precision does not mean describing every object. It means spending words on the details that change recognition, motion, composition, sound, or the ending beat.

  • Name two or three durable subject traits that remain visible during motion.
  • Describe one primary action before adding secondary motion or decoration.
  • Choose a camera move for a reason: reveal, intimacy, scale, or clarity.
  • Finish on a holdable image instead of writing only “the video ends.”

02 · Input strategy

Change the formula when the starting material changes

For text-to-video, the prompt carries the full visual brief. Define the subject, environment, action, camera relationship, lighting, visual treatment, sound, and final state. Keep the first attempt focused enough that the result can be judged. If the frame needs an exact product angle, character design, or composition, move to image-to-video instead of adding more prose in the hope that text will lock every detail.

For image-to-video, the source frame already establishes much of the subject, scene, palette, and composition. Repeating a complete visual description can introduce conflict. Concentrate on primary motion, secondary environmental motion, camera behavior, preservation boundaries, and the ending state. State what must remain recognizable—such as face shape, garment, label position, product silhouette, or the original framing—then avoid simultaneous changes to pose, camera angle, lighting, and location.

For multimodal reference work, give every upload one job and call it by its numbered token. A useful manifest might assign Image 1 to character identity, Image 2 to the room layout, Video 1 to movement rhythm, and Audio 1 to timing. Relationships belong in the prompt: explain which subject acts in which setting and which reference controls motion or sound. An asset without a clear role adds ambiguity rather than control.

  • Text to video: describe the whole shot from subject through ending.
  • Image to video: describe change and preservation, not the whole still image again.
  • Reference mode: assign each image, video, and audio file a single explicit role.
  • First-and-last frames: describe the transition that connects the two approved states.

03 · Direction

Separate subject motion from camera motion

Many weak prompts mix movement into one sentence: the person turns, fabric moves, traffic crosses, the camera orbits, and the lens zooms. Separate the hierarchy. First state what the main subject does, including direction, speed, and amplitude. Then add one or two supporting movements, such as steam rising or leaves reacting after the subject passes. Finally choose the camera behavior and say what it keeps framed.

A locked camera is useful for product clarity, performance, or subtle image animation. A slow push-in directs attention and can make a realization feel intimate. A pull-back reveals scale or isolation. A lateral track connects the viewer to movement through a space. An orbit describes form, but a wide orbit also increases the amount of unseen geometry the model must invent. Match the move to the purpose and keep it physically compatible with the subject action.

Write preservation limits beside ambitious movement. For an image-led portrait, ask for a small head turn and a gentle push rather than a fast spin with a radical angle change. For a product, keep the silhouette, material, label placement, and number of parts fixed while changing only light or camera position. The prompt should make the hierarchy obvious: one hero action, one camera idea, and a small number of environmental responses.

04 · Sequence and audio

Give longer clips a timeline and every sound a source

Wan 3.0 can produce a clip up to 30 seconds in the standard text and frame-led workflows exposed on this site. Extra time should create development, not invite a list of disconnected scenes. For a multi-shot prompt, begin with a one-sentence overall idea, then label each shot and add time ranges. A practical rhythm is establish, develop, turn, resolve, and hold. Fewer beats are better when a performance or product action needs room to read.

Keep transitions explicit. “Hard cut” is appropriate when the visual change should be immediate. A match cut needs a shared shape, motion, or framing relationship. A continuous camera move should not also contain unexplained jumps in location. Give each shot a subject action and camera position, then check that the time allocation is plausible. Five different actions will not become clear merely because they are placed inside a five-second range.

Plan sound in layers: voice, effects, ambience, and music. Put spoken words in quotation marks and specify who says them, in what language, and with what pace or emotional restraint. Name effects by their physical source—the snap of a latch, ceramic placed on stone, rain on a metal awning. Describe whether music is absent, continuous, or enters at a particular beat. If clean silence matters, direct it. “No dialogue; no music; only close room tone” is a useful audio decision.

  • Write an overall story direction before numbered shots.
  • Use time ranges only when they clarify pace and sequence.
  • Place dialogue and sound cues beside the action they accompany.
  • Reserve the final second for a stable end frame when an edit point matters.

05 · Iteration

Diagnose the first result before rewriting the prompt

Watch the first result once without stopping and ask whether the intended message reads. Then review subject identity, geometry, action, camera path, background stability, lighting continuity, sound timing, and the ending separately. Choose the highest-impact mismatch. A useful revision names that problem, changes the matching instruction, and preserves the decisions that already work.

If motion feels weak, clarify the path, speed, and visible effect of the action. If the camera fights the subject, replace multiple moves with one. If a face or product drifts, reduce pose and viewpoint change while restating the protected traits. If sound feels crowded, separate voice, effects, and music or remove a layer. If the ending cuts abruptly, add a stable final state and a short hold rather than stretching every earlier beat.

Keep the original prompt, selected model, mode, duration, aspect ratio, resolution, audio choice, references, and one-line revision note beside every useful result. This turns prompting into a controlled production process. It also prevents a team from choosing between unrelated attractive clips without knowing which instruction created the improvement.

  • Change one major variable per rerun so cause and effect remain visible.
  • Preserve successful language instead of rebuilding the entire brief.
  • Reduce conflict before adding detail; more words do not repair incompatible directions.
  • Review the output at the size and orientation where it will actually be used.

Prompt library

Copy a complete brief, then make it yours

Replace bracketed details, remove directions that do not serve the shot, and confirm the selected mode before generation.

Ads

Ad prompts with one message and a clean end card

Keep the claim outside the generated footage when accuracy matters; use the video to make one product benefit visually legible.

01

Ads

Mechanical product reveal

A single-shot demonstration that protects product geometry and creates a clean space for added copy.

Single continuous product shot. A compact matte-black travel coffee grinder stands centered on pale limestone. The top cap rotates one quarter-turn, lifts smoothly, and reveals the brushed-steel burr assembly; no extra parts appear. Begin with a low three-quarter medium shot, then make a slow clockwise orbit of no more than 25 degrees while keeping the grinder centered. Hard side light defines the knurled metal and soft fill preserves shadow detail. Quiet studio room tone with one precise mechanical click as the cap releases; no music and no voice. Preserve the cylindrical silhouette, logo-free surface, and exact number of components. End on the open grinder with clear negative space on the left and hold for one second.

Opens with this prompt ready to editUse this prompt on wan-3.ai
02

Ads

Service benefit in one moment

Show the benefit through a small before-and-after action rather than relying on generated text.

Create a 9:16 social ad about fast grocery delivery without showing any brand or written claim. In a bright apartment kitchen at early evening, a tired parent opens an almost empty refrigerator, pauses, then hears the doorbell. Hard cut to the same person placing fresh vegetables and bread on the counter with visible relief. Use natural handheld framing with restrained movement, warm practical light, and believable domestic detail. Sound: refrigerator hum, distant city room tone, one clear doorbell, paper bag rustle; light upbeat percussion begins only after the cut. Keep wardrobe, kitchen layout, and time of day consistent. Finish on the full counter with clean upper-frame space for a caption added later.

Opens with this prompt ready to editUse this prompt on wan-3.ai

UGC

UGC prompts that feel observed, not over-produced

Use a simple performance, plausible phone framing, and restrained background motion so the message remains readable.

03

UGC

First-use reaction

A compact reaction clip with controlled dialogue and no invented on-screen interface.

Vertical phone-style UGC video in a quiet home office. A creator in a charcoal sweatshirt sits at a desk, looks just beside the lens at an unseen laptop, raises their eyebrows, and turns to camera with a surprised half-smile. Locked chest-up framing with a slight natural handheld drift; do not zoom. Soft window light from camera left, neutral wall, one plant moving slightly from an air vent. The creator says in conversational American English, “Okay, that took less setup than I expected,” with a short pause after “Okay.” Keep lip movement aligned with the line. No captions, logos, background music, or extra people. End as the creator looks back to the laptop and gives one small approving nod.

Opens with this prompt ready to editUse this prompt on wan-3.ai
04

UGC

Three-beat mini review

A timed structure makes a short review feel spontaneous while keeping each beat distinct.

Create a 12-second vertical creator review in one apartment location. Shot 1 [0–4s]: handheld medium close-up; the creator holds an unbranded insulated bottle and says, “I wanted something that would not leak in my work bag.” Shot 2 [4–8s]: hard cut to an overhead close-up as they close the lid, turn the bottle upside down over a dry towel, and wait; only cap clicks and room tone. Shot 3 [8–12s]: hard cut back to the original framing; they show the dry towel and say, “So far, so good,” with a restrained smile. Keep the bottle color, clothing, hands, room, and daylight consistent. No subtitles, beauty filter, or music. Hold the final reaction for half a second.

Opens with this prompt ready to editUse this prompt on wan-3.ai

Product

Product prompts with geometry and material boundaries

Name the details that establish recognition, then let light and one controlled camera move do the visual work.

05

Product

Material surface study

Useful for a quiet premium product shot where texture matters more than spectacle.

Single shot of an unbranded dark-green ceramic fragrance bottle on wet black stone. Preserve the rectangular bottle geometry, short cylindrical cap, blank label area, and one bottle only. Start on a macro detail of condensation along the glass edge, then pull back slowly to a centered three-quarter product view. A narrow cool highlight travels across the glass while a warm reflected line appears on the cap; droplets move naturally downward but the bottle remains still. Deep forest palette, crisp material detail, shallow depth of field, no flowers, smoke, hands, text, logo, or duplicate container. Sound: isolated water drops and very low room tone; no music. End with the whole bottle sharp and negative space above it.

Opens with this prompt ready to editUse this prompt on wan-3.ai
06

Product

Animate an approved product frame

Upload the approved still first; this prompt directs movement without redescribing or replacing the composition.

Use the uploaded image as the exact opening composition. Keep the product silhouette, packaging color, label position, typography shapes, camera angle, surface, and background unchanged. Add only three motions: a soft band of window light travels slowly from left to right across the package, one small reflection shifts along the front edge, and background dust catches the light briefly. Locked camera with no pan, orbit, zoom, crop, or lens change. No new props, hands, text, logos, liquids, or duplicate products. Sound: quiet studio ambience with one subtle paper-texture rustle; no voice and no music. End with the lighting slightly warmer than the opening while every object remains in its original position.

Opens with this prompt ready to editUse this prompt on wan-3.ai

Multi-shot

Multi-shot prompts with explicit timing

State the overall idea once, then give every shot a time range, action, camera relationship, and transition.

07

Multi-shot

From raw material to finished object

A three-shot craft story with one material, one location, and a match-cut transition.

Overall direction: a restrained 15-second process film about a ceramic cup moving from raw clay to a finished morning ritual, with consistent warm workshop light. Shot 1 [0–5s]: macro side view of two hands centering wet clay on a spinning wheel; fixed camera, visible water and clay texture, wheel motor and hand friction only. Shot 2 [5–10s]: match cut on the circular rim to the fired cup being lifted from a kiln shelf with padded tongs; slow pull-back, low fire crackle and room ambience. Shot 3 [10–15s]: hard cut to the same cup on a wooden breakfast table as tea is poured; gentle push-in, steam rises, a spoon touches ceramic once. Keep the cup proportions and speckled cream glaze consistent after firing. No dialogue, music, text, logo, extra cups, or abrupt camera shake. End on the steam crossing morning light.

Opens with this prompt ready to editUse this prompt on wan-3.ai
08

Multi-shot

A small narrative turn

A short character arc that uses sound and framing to move from pressure to relief.

Overall direction: an 18-second urban micro-story about a bicycle courier finding a quiet pause after rain. Shot 1 [0–6s]: wide lateral tracking shot as the courier rides through wet evening traffic, shoulders tense, tires spraying a small amount of water; layered street ambience, no music. Shot 2 [6–12s]: hard cut to a sheltered arcade; medium fixed frame as the courier stops, removes the helmet, and notices a shaft of sunset between buildings; traffic becomes muffled and one distant bird is audible. Shot 3 [12–18s]: slow push from medium shot to close-up as the courier exhales and smiles slightly; one soft sustained synth note enters. Preserve the yellow jacket, black bicycle, wet pavement, and direction of light across every shot. End on the reflected orange sky in the courier’s eyes and hold.

Opens with this prompt ready to editUse this prompt on wan-3.ai

Audio

Audio prompts with clean layer assignments

Put speech, physical effects, ambience, and music in separate clauses, then say when each layer begins or stops.

09

Audio

Dialogue with a physical cue

The action creates a natural timing marker for one short line and a quiet reaction.

Single continuous medium two-shot in a small bakery before opening. A baker slides a warm loaf onto the counter; a colleague leans closer, waits for the paper bag to stop rustling, then says in calm British English, “That is the one for the window.” The baker looks at the loaf, gives a brief satisfied nod, and turns it so the scored crust faces the street. Locked camera at counter height, soft dawn light through the front glass, realistic flour and paper textures. Audio layers: quiet refrigerator hum and distant street ambience throughout; wooden board contact at the start; paper rustle before the line; no overlapping speech; no music. Keep both faces, aprons, loaf shape, and counter layout consistent. End on the loaf centered between them for one second.

Opens with this prompt ready to editUse this prompt on wan-3.ai
10

Audio

Sound-led object rhythm

A product sequence whose cuts follow real object sounds instead of an unspecified energetic soundtrack.

Create a 10-second sound-led desk organization film with no people visible above the wrists. Shot 1 [0–3s]: overhead close-up as a notebook lands squarely on a walnut desk; cut exactly on the soft paper-and-wood impact. Shot 2 [3–6s]: side macro view as a metal pen clicks once and slides into the notebook loop; cut on the click. Shot 3 [6–10s]: front three-quarter view as a small desk lamp switches on and warm light spreads across the arrangement; hold after the switch sound. Keep the navy notebook, silver pen, walnut grain, and object positions consistent. Audio: isolated natural impacts with clean room tone; a minimal two-note bass pulse follows the final lamp click only. No dialogue, captions, logos, extra stationery, rapid camera motion, or duplicate hands.

Opens with this prompt ready to editUse this prompt on wan-3.ai

Practical answers

Questions creators ask before a run

How long should a Wan 3.0 prompt be?+

Use enough detail to resolve the important decisions, not every possible detail. A focused single shot may need one compact paragraph. A longer multi-shot video benefits from an overall direction plus numbered, timed shots. If two clauses compete for subject, location, camera, or mood, shorten or split the idea before adding more adjectives.

Should I put camera terms in every prompt?+

Only when camera behavior changes how the idea communicates. A locked frame can be the best direction for a product demonstration or restrained portrait. When you use a move, name one primary move and what it keeps framed. Stacking orbit, crane, zoom, pan, and handheld shake makes a short shot harder to direct.

Can I use the same prompt for text-to-video and image-to-video?+

Use the same creative goal, but change the emphasis. Text-to-video needs the full subject, scene, action, camera, look, and sound brief. Image-to-video should trust the source frame for appearance and composition, then focus on movement, camera restraint, preservation boundaries, and the final state.

How do references work inside a Wan 3.0 prompt?+

Upload the assets in the generator, then refer to them by their numbered image, video, or audio token. Give each reference a distinct job and describe the relationship between them. The current Wan 3.0 workspace accepts up to 10 images, 5 videos, and 5 audio clips in reference mode, subject to file and duration limits shown in the interface.

Why does a detailed prompt still produce an unexpected result?+

Detail cannot remove model variability or repair contradictory direction. Check whether the subject action, camera move, duration, reference roles, and ending can all coexist. Then revise the largest mismatch first. Keep the model, settings, references, and successful clauses fixed so the next result tests one identifiable change.

Keep building the workflow

Product references

Model limits and prompt structures were checked against the current wan-3.ai generator contract and the primary documentation below. Creative recommendations are editorial guidance, not a guarantee of a specific output.