Explore indexBack to Terms

Multimodal Reference

Multimodal Reference is a conditioning technique in video generation that incorporates multiple media types—such as images, video clips, audio tracks, and 3D greybox models—into the model. It extracts identity, pose, style, and temporal features via cross-attention to maintain precise control over generated scenes.

No public content is connected to this entity yet.