멀티모달 생성
하나의 모델이 말하고, 그리고, 움직이고, 심지어 3차원 세계를 모델링한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
정의
Multimodal generation uses a unified model to handle and produce several modalities — text to image, video, audio and 3D structure. The central idea is to cut every modality into discrete tokens or continuous latents and hand them to the same sequence-modelling or diffusion machinery: video is split into spatiotemporal patches, speech is encoded into acoustic tokens, and 3D scenes are represented as radiance fields or Gaussian splats.
직관적 이해
Text, images, video and sound are different carriers sampling the same underlying information. Video is just images with a time axis, audio is a waveform sampled at a high rate, and 3D is multi-view imagery constrained to be geometrically consistent. Once a model learns to predict the next unit in a unified representation, the walls between modalities become a translation problem rather than three incompatible methodologies.
Key milestones in generative AI: from VAE and GAN, through diffusion and latent diffusion, to 3D and video generation
Relative maturity of four generation modalities (qualitative: quality, cross-step consistency, controllability, low cost, data availability; higher is better)
- Image generation
- Video generation
- Speech synthesis
- 3D generation
작동 원리
- 01
Unify tokenisation across modalities
Text is cut into subword tokens, images into patches, video into spatiotemporal patches, audio into discrete acoustic codes. Once everything is a unit, the remaining task is structurally identical to language modelling.
- 02
Spatiotemporal patches: video’s extra dimension
Video adds a time axis to images. The common approach groups patches at the same location across several frames into a tube-shaped patch, or encodes keyframes and interpolates the ones between, modelling motion within a manageable compute budget.
- 03
Dedicated representations for speech and 3D
Speech synthesis first converts text into intermediate acoustic features with an acoustic model, then restores the waveform with a vocoder; 3D uses neural radiance fields (NeRF) or 3D Gaussian splatting to represent a scene, with diffusion models generating or editing those representations.
- 04
Toward a unified model
The trend shifts from one model per modality to a single backbone for all. The cost is that modalities differ sharply in compute demand, sampling rate and evaluation, so unified training requires careful balancing of data mixes and loss weights.
Video is split into a grid of frames × spatial patches; each column forms a tubelet spanning time, the basic unit for modelling motion in text-to-video
응용 분야
- Text-to-video, video editing, and animation or VFX assistance
- Speech synthesis, voice cloning and music generation
- 3D content creation: reconstructing and generating objects, scenes and digital humans
- Unified assistants: one model for both understanding and generation, enabling cross-modal dialogue
흔한 오해
- The hard part of video generation is not per-frame quality but temporal consistency: flicker, sudden object changes and physically implausible motion are the main failure modes.
- Evaluating 3D generation is far harder than 2D: 2D metrics cannot measure geometric consistency, and ground truth is often hard to obtain.
- Unifying modalities does not simply add capabilities. Merging backbones can sacrifice single-modality peak performance, and negative transfer between modalities is real.
핵심 용어
- Spatiotemporal patch
- A local unit spanning frames in video, used to model motion
- Vocoder
- The component that turns acoustic features back into a waveform
- NeRF
- A neural network representing a scene’s radiance field for novel-view synthesis
- 3D Gaussian Splatting
- Representing a scene with many 3D Gaussian ellipsoids for fast rendering