Reference-to-video is not a contest to upload the most material. It is a production system in which every image, clip, audio file, document, or link needs an explicit job.
Wan 3.0 can work with multiple media types and, in its documented reference workflow, supports up to 10 images, 5 videos, and 5 audio clips, as well as document or web-page input. That capacity is useful only when the assets agree about the scene you are asking the model to make.
This guide shows how to build a reference manifest, choose the smallest useful asset set, write a prompt that names relationships, and review whether the result used each source correctly. Use it with the reference-to-video workspace.
Decide whether you need references at all
Use references when the output must carry forward approved information:
- a person, character, creature, or product identity;
- wardrobe, material, packaging, or prop detail;
- a location, spatial layout, interface, or graphic system;
- a motion pattern, performance rhythm, or camera behavior;
- a voice, ambience, sound texture, or musical pulse;
- facts or structure from a supported document or public page.
Use text-to-video when the idea can be invented freely. Use image-to-video when one opening frame—or an opening and ending frame—already defines the shot. Reference mode is most valuable when several different pieces of source material must cooperate.
Build a reference manifest before uploading
Create a small table in your project notes:
| Label | Source | Role | Must preserve | May ignore |
|---|---|---|---|---|
| Image 1 | lead portrait | identity | face shape, hair, navy coat | source background |
| Image 2 | station interior | location | green tile, arch spacing, clock position | people in source |
| Video 1 | walking reference | performance | measured pace, shoulder rhythm | performer identity, original setting |
| Audio 1 | field recording | ambience | train brake texture, platform reverb | distant announcement words |
This manifest separates content from role. Video 1 may show a different person in a different place; you are using only its motion. If the prompt does not say that, the model has to infer which parts matter.
Apply the one-job rule
Give each reference one primary responsibility. An image can naturally contribute several details, but your direction should establish a priority.
Good roles include:
- identity anchor;
- wardrobe or material anchor;
- environment layout;
- composition reference;
- motion rhythm;
- camera path;
- voice or ambience;
- factual source.
Vague roles such as “inspiration,” “make it like this,” or “use everything” are difficult to verify. You cannot tell whether the result failed if success was never defined.
Remove redundant and contradictory assets
Before submission, compare every pair of sources.
Identity conflict
Two portraits may show different ages, hairlines, makeup, face proportions, or wardrobe. Decide whether they represent the same character from different angles or different characters. State that explicitly.
Spatial conflict
Two location images may disagree about door position, window count, ceiling height, or light direction. Pick one layout authority and use the other only for a named detail.
Motion conflict
A slow, weighted performance reference and a fast handheld camera reference can pull the scene toward incompatible rhythms. Assign one to subject movement and the other to camera—or remove one.
Sound conflict
Dialogue, ambience, music, and effects can all compete. Specify which audio is content, which is rhythm, and which should not be reproduced literally.
Style conflict
Photorealistic identity art, graphic animation, glossy product lighting, and documentary footage do not automatically form one coherent visual system. Name the final treatment and identify which source controls it.
If two assets do the same job, keep the clearer one. If they disagree on a critical fact, resolve the disagreement before generating.
Use stable reference labels
In the wan-3.ai workflow, insert references with the labels provided by the interface, such as @image1, @video1, and @audio1. Keep each label beside the noun and role it controls.
Weak:
Use the images and videos to make them walk through the station with the same feel.
Stronger:
The person from @image1 walks through the station from @image2. Preserve @image1’s face, hair, and navy coat; use @image2 only for the green-tile arches, clock position, and cool overhead light. Apply the measured walking pace and shoulder rhythm from @video1, but do not copy its performer, clothing, or location. Use @audio1 only for platform ambience and train-brake texture; do not reproduce intelligible announcement words.
The second version gives every source a boundary. Boundaries prevent an asset from influencing parts of the output it was never meant to control.
Write relationships, not inventory
A reference prompt should describe how sources interact in the new scene.
Use this structure:
[Identity source] performs [action informed by motion source] inside [environment source], while [audio source] controls [sound role]. Preserve [critical details]. Ignore [irrelevant source details]. Camera [movement and framing]. End on [final state].
Complete example:
The person from @image1 enters the station from @image2 and pauses beneath the clock. Preserve @image1’s facial proportions, short dark hair, navy wool coat, and silver lapel pin. Use @image2 as the layout authority for the green-tile arches, central clock, platform edge, and cool overhead lighting; remove the people visible in the source. Apply only the measured walking pace and slight shoulder weight from @video1, not its performer or outdoor background. The camera tracks beside the character at chest height, then stops as the character looks up at the clock. @audio1 supplies low station reverb and the texture of a distant train braking; no intelligible announcement and no music. End in a medium profile with the clock visible above the character, and hold.
This prompt is longer than a text-only idea because it functions as an asset contract.
Choose frames or references before upload
The current Wan 3.0 workflow treats first/last-frame control and multimodal reference input as separate paths. Choose based on the most important constraint:
- choose frame mode when an approved opening or ending composition is non-negotiable;
- choose reference mode when identity, location, movement, sound, document, or web information needs to be combined;
- do not upload a boundary frame and a reference pack with the expectation that every control will be mixed into one request.
The interface validates compatible combinations and current limits. Resolve the mode before writing a prompt so you do not build direction around inputs that cannot be submitted together.
Use documents and web pages as information sources
A document or public page is different from a visual identity reference. Treat it as a source of facts, sequence, terminology, or explanation structure.
Before using it:
- define the audience and the one takeaway the video should communicate;
- extract the exact names, numbers, steps, and claims that must remain accurate;
- decide which information should be spoken, shown, or left out;
- remove sensitive, confidential, outdated, or irrelevant material;
- plan a manual fact check against the original source after generation.
Do not ask the model to “make a video about the entire report.” Select a section and a communication goal.
Example direction:
Use the uploaded report only for the three-stage maintenance sequence and the names of the components. Explain the sequence to new operators in chronological order. Do not invent performance figures or safety thresholds. Show each component name as a simple lower-third only when that component is discussed. End with a reminder to follow the approved maintenance manual.
Generated text, charts, numbers, and narration still require review. A source document improves grounding; it does not replace editorial responsibility.
Design audio references with boundaries
An audio file can serve different jobs:
- voice identity or delivery reference;
- timing and rhythm;
- ambience;
- sound-effect texture;
- music mood.
Name the job and what should not transfer.
Use @audio1 for the speaker’s calm cadence and warm vocal tone, but use the new scripted line exactly. Do not reproduce the background music or room noise from @audio1.
Or:
Use @audio2 only as a rhythmic guide for movement accents. Generate no vocals and no recognizable melody from the reference.
Clear boundaries make both creative review and rights review easier.
Prune the pack before generation
Run this test for every reference:
- Can I state its role in one sentence?
- Is that role already covered by a clearer source?
- Does it contradict the source that controls identity, space, motion, or sound?
- Will the requested action reveal information this source does not provide?
- Do I have permission to use it for this project?
Remove any asset that fails the first question. Resolve any asset that fails the third or fifth.
Review reference fidelity by role
Do not ask only, “Does it look like the references?” Check each contract separately.
| Role | Review question | Where to inspect |
|---|---|---|
| Identity | Are defining facial, wardrobe, product, or character traits stable? | profile changes, close-ups, contact, fast motion |
| Location | Does the spatial layout remain readable? | entries, exits, camera turns, background lines |
| Motion | Does the pace and weight match without copying irrelevant content? | start, peak action, recovery |
| Camera | Is the intended path followed without breaking subject visibility? | acceleration, orbit, reframing |
| Audio | Did the intended voice, rhythm, ambience, or effect transfer at the right time? | dialogue, transient effects, scene changes |
| Information | Are names, numbers, steps, and visible text accurate? | captions, narration, charts, final frame |
Write a time code for the first failure. Then revise the role boundary, remove the conflicting source, or reduce the action that caused the breakdown.
Reference pack checklist
- Every source has one primary job.
- One asset is named as the authority for each critical identity or layout decision.
- Redundant and contradictory sources are removed.
- Interface labels are used consistently in the prompt.
- The prompt states both what to preserve and what to ignore.
- Frame mode and reference mode are not mixed accidentally.
- Document and web facts have a manual verification plan.
- Audio roles separate voice, rhythm, ambience, effects, and music.
- Rights and consent are confirmed for every uploaded source.
- Review notes evaluate each asset according to its assigned role.
When the manifest is clean, open Wan 3.0 reference-to-video. If one still image already contains everything the shot needs, the simpler image-to-video workflow may give you a clearer first test.


