Génération multimodale
Un seul modèle qui apprend à parler, dessiner, bouger — et même à modéliser le monde 3D
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
DÉFINITION
Multimodal generation uses a unified model to handle and produce several modalities — text to image, video, audio and 3D structure. The central idea is to cut every modality into discrete tokens or continuous latents and hand them to the same sequence-modelling or diffusion machinery: video is split into spatiotemporal patches, speech is encoded into acoustic tokens, and 3D scenes are represented as radiance fields or Gaussian splats.
Intuition
Text, images, video and sound are different carriers sampling the same underlying information. Video is just images with a time axis, audio is a waveform sampled at a high rate, and 3D is multi-view imagery constrained to be geometrically consistent. Once a model learns to predict the next unit in a unified representation, the walls between modalities become a translation problem rather than three incompatible methodologies.
Key milestones in generative AI: from VAE and GAN, through diffusion and latent diffusion, to 3D and video generation
Relative maturity of four generation modalities (qualitative: quality, cross-step consistency, controllability, low cost, data availability; higher is better)
- Image generation
- Video generation
- Speech synthesis
- 3D generation
Fonctionnement
- 01
Unify tokenisation across modalities
Text is cut into subword tokens, images into patches, video into spatiotemporal patches, audio into discrete acoustic codes. Once everything is a unit, the remaining task is structurally identical to language modelling.
- 02
Spatiotemporal patches: video’s extra dimension
Video adds a time axis to images. The common approach groups patches at the same location across several frames into a tube-shaped patch, or encodes keyframes and interpolates the ones between, modelling motion within a manageable compute budget.
- 03
Dedicated representations for speech and 3D
Speech synthesis first converts text into intermediate acoustic features with an acoustic model, then restores the waveform with a vocoder; 3D uses neural radiance fields (NeRF) or 3D Gaussian splatting to represent a scene, with diffusion models generating or editing those representations.
- 04
Toward a unified model
The trend shifts from one model per modality to a single backbone for all. The cost is that modalities differ sharply in compute demand, sampling rate and evaluation, so unified training requires careful balancing of data mixes and loss weights.
Video is split into a grid of frames × spatial patches; each column forms a tubelet spanning time, the basic unit for modelling motion in text-to-video
Où c'est utilisé
- Text-to-video, video editing, and animation or VFX assistance
- Speech synthesis, voice cloning and music generation
- 3D content creation: reconstructing and generating objects, scenes and digital humans
- Unified assistants: one model for both understanding and generation, enabling cross-modal dialogue
Idées fausses courantes
- The hard part of video generation is not per-frame quality but temporal consistency: flicker, sudden object changes and physically implausible motion are the main failure modes.
- Evaluating 3D generation is far harder than 2D: 2D metrics cannot measure geometric consistency, and ground truth is often hard to obtain.
- Unifying modalities does not simply add capabilities. Merging backbones can sacrifice single-modality peak performance, and negative transfer between modalities is real.
Termes clés
- Spatiotemporal patch
- A local unit spanning frames in video, used to model motion
- Vocoder
- The component that turns acoustic features back into a waveform
- NeRF
- A neural network representing a scene’s radiance field for novel-view synthesis
- 3D Gaussian Splatting
- Representing a scene with many 3D Gaussian ellipsoids for fast rendering