Direct Wan 3.0 with images, video, and audio

Wan 3.0 Reference to Video AI Generator

Combine up to 10 images, 5 video clips, and 5 audio clips in one Wan 3.0 production brief. Use @ mentions to tell the model which reference should guide the subject, motion, camera, timing, or sound.

Standard Wan 3.0 opens by default. You can still switch models, and the available reference types and limits will follow the current selection.

Images, video clips, and audio arranged as references for an AI video brief

Assign Every Wan 3.0 Reference One Production Decision

Treat the reference set as a directed shot plan. Decide which asset controls identity, environment, motion, camera rhythm, or sound before Wan 3.0 combines them.

Images anchor the look

Use up to 10 images to establish a recognizable subject, product, character, palette, environment, or composition.

Clips describe movement

Use up to 5 video clips when performance, camera behavior, pacing, or a transition is easier to demonstrate than describe.

Audio sets timing

Use up to 5 audio clips to establish rhythm, dialogue cadence, ambience, or a cue that should meet a visible moment.

The prompt connects them

Mention each upload with @image, @video, or @audio and state exactly which creative decision it should control.

Build a Reference Set Wan 3.0 Can Follow

Keep only assets that change the intended scene, connect them to explicit prompt instructions, then review whether every assigned role survived.

Include a file only when it defines identity, setting, motion, camera language, timing, or sound that would otherwise be ambiguous.

A Wan 3.0 Multimodal Reference Workflow

Coordinate visual, motion, and audio references, then create a 2–30 second Wan 3.0 video from one directed brief.

Mixed-media references

Bring images, video clips, and audio together to guide one new scene.

Explicit asset mentions

Use @image, @video, and @audio tokens to connect prompt directions to uploaded files.

Model-aware settings

Wan 3.0 supports 2–30 seconds, 480P, 720P, 1080P, and audio; controls update if you switch models.

Frame-based alternative

When a shot needs fixed endpoints, choose first and last frames instead of multimodal references.

Visible task progress

Follow generation status and reopen completed results from My Creations.

Downloadable result

Preview the completed clip and download the returned video when it is ready.

Reference-to-Video AI Questions

What to know before generating a new video from mixed reference assets.

What is a reference-to-video AI generator?

A reference-to-video AI generator creates a new clip from a prompt plus source assets such as images, video clips, or audio. Wan 3.0 can accept up to 10 images, 5 videos, and 5 audio clips, with each reference assigned to identity, style, motion, camera language, rhythm, or sound.


How is reference to video different from image to video?

Image to video usually begins with one image or a frame pair that anchors the shot. Reference to video can assign separate roles to multiple supported assets, including images, motion clips, and audio cues.


Can Wan 3.0 use images, videos, and audio in one request?

Yes. Wan 3.0 can combine up to 10 images, 5 video clips, and 5 audio clips in its multimodal reference mode. Give every upload a distinct role and mention it with the matching @ token. A focused set with clear assignments is easier to direct than filling every available slot.


How do I tell Wan 3.0 which reference file to use?

Upload the assets, type @ in the prompt, and select the matching reference token. Then state the role of that asset, such as subject from @image1, camera movement from @video1, or rhythm from @audio1. Place the instruction close to the matching token and avoid assigning conflicting references to the same decision unless you explain their priority.


When should Wan 3.0 use first and last frames instead?

Choose first and last frames when the shot needs fixed visual endpoints. Choose multimodal references when separate images, motion clips, or audio cues should influence the scene.


Why is Wan 3.0 selected for Reference to Video?

Wan 3.0 is selected by default because it supports image, video, and audio references together. You can switch to another model when its input types, duration, resolution, sound controls, or cost fit the project better.


How should I prepare a Wan 3.0 multimodal reference brief?

Start with the new scene you want Wan 3.0 to create, then assign one clear job to every uploaded asset. Identify which image controls the subject or visual style, which video demonstrates motion or camera language, and which audio supplies dialogue, rhythm, ambience, or an effect. Put the visible action in chronological order and finish with the intended ending state. Remove any reference whose role cannot be explained in one sentence.


How should I divide Wan 3.0's 10 image reference slots?

Wan 3.0 accepts up to 10 reference images, but every slot should earn its place. Use the smallest set that covers identity, product details, wardrobe, environment, composition, or visual treatment, and name each role beside its @ token. If one image controls subject identity and another controls lighting or palette, say so explicitly. Remove near-duplicates and resolve conflicting colors, shapes, or character details before submitting.


How can a Wan 3.0 video reference control motion without replacing identity?

Attach the clip with its @ token and name only the motion property to borrow, such as camera speed, a subject gesture, blocking, transition timing, or action rhythm. Define the new subject and environment with separate image references or prompt text. Wan 3.0 accepts up to 5 reference videos, so avoid assigning several clips to the same motion decision unless you also state which one has priority.


How do Wan 3.0 audio references line up with visible action?

Wan 3.0 accepts up to 5 audio references. Use an @ token for each clip, state whether it controls dialogue cadence, musical rhythm, ambience, or a sound effect, and connect important beats to visible actions in the prompt timeline. Keep the planned action inside the selected 2–30 second duration. Review the result with sound on to check synchronization, then with sound off to confirm the visual story still reads clearly.


What should I review against each @ reference after Wan 3.0 generates?

Turn the prompt into a checklist: subject identity from the assigned image, motion from the assigned video, and rhythm or sound from the assigned audio. Then inspect continuity, camera movement, background stability, product or facial details, ending state, and audio timing. If one source did not perform its stated role, change that asset or its nearby instruction before altering unrelated settings.


How should a team version Wan 3.0 reference sets?

Record the Wan model, prompt, duration, resolution, sound setting, uploaded filenames, @ token mapping, and the role assigned to every reference. Label each new result by the single variable being tested, such as a replacement identity image or a different motion clip. Keep the strongest output in My Creations and compare versions with the same identity, motion, camera, and audio checklist.


What should I do when Wan 3.0 references conflict?

Rank subject identity, product attributes, action, camera, style, and sound by the project goal, then give every asset one primary job. If two images disagree on color or shape, state beside their @ tokens which one controls the result. If a motion clip conflicts with the still composition, validate the subject in a simpler shot first. Removing a reference with no clear role is usually more effective than adding more restrictions.


How should I prepare a Wan 3.0 multimodal reference brief?

Start with the new scene you want Wan 3.0 to create, then assign one clear job to every uploaded asset. Identify which image controls the subject or visual style, which video demonstrates motion or camera language, and which audio supplies dialogue, rhythm, ambience, or an effect. Put the visible action in chronological order and finish with the intended ending state. Remove any reference whose role cannot be explained in one sentence.


How should I divide Wan 3.0's 10 image reference slots?

Wan 3.0 accepts up to 10 reference images, but every slot should earn its place. Use the smallest set that covers identity, product details, wardrobe, environment, composition, or visual treatment, and name each role beside its @ token. If one image controls subject identity and another controls lighting or palette, say so explicitly. Remove near-duplicates and resolve conflicting colors, shapes, or character details before submitting.


How can a Wan 3.0 video reference control motion without replacing identity?

Attach the clip with its @ token and name only the motion property to borrow, such as camera speed, a subject gesture, blocking, transition timing, or action rhythm. Define the new subject and environment with separate image references or prompt text. Wan 3.0 accepts up to 5 reference videos, so avoid assigning several clips to the same motion decision unless you also state which one has priority.


How do Wan 3.0 audio references line up with visible action?

Wan 3.0 accepts up to 5 audio references. Use an @ token for each clip, state whether it controls dialogue cadence, musical rhythm, ambience, or a sound effect, and connect important beats to visible actions in the prompt timeline. Keep the planned action inside the selected 2–30 second duration. Review the result with sound on to check synchronization, then with sound off to confirm the visual story still reads clearly.


What should I review against each @ reference after Wan 3.0 generates?

Turn the prompt into a checklist: subject identity from the assigned image, motion from the assigned video, and rhythm or sound from the assigned audio. Then inspect continuity, camera movement, background stability, product or facial details, ending state, and audio timing. If one source did not perform its stated role, change that asset or its nearby instruction before altering unrelated settings.


How should a team version Wan 3.0 reference sets?

Record the Wan model, prompt, duration, resolution, sound setting, uploaded filenames, @ token mapping, and the role assigned to every reference. Label each new result by the single variable being tested, such as a replacement identity image or a different motion clip. Keep the strongest output in My Creations and compare versions with the same identity, motion, camera, and audio checklist.


What should I do when Wan 3.0 references conflict?

Rank subject identity, product attributes, action, camera, style, and sound by the project goal, then give every asset one primary job. If two images disagree on color or shape, state beside their @ tokens which one controls the result. If a motion clip conflicts with the still composition, validate the subject in a simpler shot first. Removing a reference with no clear role is usually more effective than adding more restrictions.


How should I prepare a Wan 3.0 multimodal reference brief?

Start with the new scene you want Wan 3.0 to create, then assign one clear job to every uploaded asset. Identify which image controls the subject or visual style, which video demonstrates motion or camera language, and which audio supplies dialogue, rhythm, ambience, or an effect. Put the visible action in chronological order and finish with the intended ending state. Remove any reference whose role cannot be explained in one sentence.


How should I divide Wan 3.0's 10 image reference slots?

Wan 3.0 accepts up to 10 reference images, but every slot should earn its place. Use the smallest set that covers identity, product details, wardrobe, environment, composition, or visual treatment, and name each role beside its @ token. If one image controls subject identity and another controls lighting or palette, say so explicitly. Remove near-duplicates and resolve conflicting colors, shapes, or character details before submitting.


How can a Wan 3.0 video reference control motion without replacing identity?

Attach the clip with its @ token and name only the motion property to borrow, such as camera speed, a subject gesture, blocking, transition timing, or action rhythm. Define the new subject and environment with separate image references or prompt text. Wan 3.0 accepts up to 5 reference videos, so avoid assigning several clips to the same motion decision unless you also state which one has priority.


How do Wan 3.0 audio references line up with visible action?

Wan 3.0 accepts up to 5 audio references. Use an @ token for each clip, state whether it controls dialogue cadence, musical rhythm, ambience, or a sound effect, and connect important beats to visible actions in the prompt timeline. Keep the planned action inside the selected 2–30 second duration. Review the result with sound on to check synchronization, then with sound off to confirm the visual story still reads clearly.


What should I review against each @ reference after Wan 3.0 generates?

Turn the prompt into a checklist: subject identity from the assigned image, motion from the assigned video, and rhythm or sound from the assigned audio. Then inspect continuity, camera movement, background stability, product or facial details, ending state, and audio timing. If one source did not perform its stated role, change that asset or its nearby instruction before altering unrelated settings.


How should a team version Wan 3.0 reference sets?

Record the Wan model, prompt, duration, resolution, sound setting, uploaded filenames, @ token mapping, and the role assigned to every reference. Label each new result by the single variable being tested, such as a replacement identity image or a different motion clip. Keep the strongest output in My Creations and compare versions with the same identity, motion, camera, and audio checklist.


What should I do when Wan 3.0 references conflict?

Rank subject identity, product attributes, action, camera, style, and sound by the project goal, then give every asset one primary job. If two images disagree on color or shape, state beside their @ tokens which one controls the result. If a motion clip conflicts with the still composition, validate the subject in a simpler shot first. Removing a reference with no clear role is usually more effective than adding more restrictions.


Direct a Wan 3.0 Scene with Your References

Keep the assets that matter, assign each one a clear role, and turn the set into one focused Wan 3.0 video brief.