Zum Inhalt springen
KI-Atlas

Multimodale Generierung

Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt

07 Generative KIFortgeschritten6. Eintrag in diesem Bereich

Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.

DEFINITION

Multimodal generation uses a unified model to handle and produce several modalities — text to image, video, audio and 3D structure. The central idea is to cut every modality into discrete tokens or continuous latents and hand them to the same sequence-modelling or diffusion machinery: video is split into spatiotemporal patches, speech is encoded into acoustic tokens, and 3D scenes are represented as radiance fields or Gaussian splats.

Intuition

Text, images, video and sound are different carriers sampling the same underlying information. Video is just images with a time axis, audio is a waveform sampled at a high rate, and 3D is multi-view imagery constrained to be geometrically consistent. Once a model learns to predict the next unit in a unified representation, the walls between modalities become a translation problem rather than three incompatible methodologies.

Abb. 1

Key milestones in generative AI: from VAE and GAN, through diffusion and latent diffusion, to 3D and video generation

2013Variational autoencoderYields a samplable latent space, giving generation a probabilistic basis2014Generative adversarial networkReplaces likelihood with an adversarial game and puts sample quality centre stage2015Diffusion theoryNon-equilibrium thermodynamics gives the add-noise-then-reverse framework2020Denoising diffusion (DDPM)Makes diffusion practical, matching and surpassing GANs in quality2021CLIP & DALL·EAligns text and images in a shared space, opening text-to-image2022Latent diffusionStable Diffusion moves diffusion into an efficient latent space and popularises it20233D Gaussian splattingA fast, differentiable 3D scene representation that advances 3D generation2024Text-to-videoDiffusion extends to spatiotemporal patches to generate coherent video
Abb. 2

Relative maturity of four generation modalities (qualitative: quality, cross-step consistency, controllability, low cost, data availability; higher is better)

  • Image generation
  • Video generation
  • Speech synthesis
  • 3D generation

Funktionsweise

  1. 01

    Unify tokenisation across modalities

    Text is cut into subword tokens, images into patches, video into spatiotemporal patches, audio into discrete acoustic codes. Once everything is a unit, the remaining task is structurally identical to language modelling.

  2. 02

    Spatiotemporal patches: video’s extra dimension

    Video adds a time axis to images. The common approach groups patches at the same location across several frames into a tube-shaped patch, or encodes keyframes and interpolates the ones between, modelling motion within a manageable compute budget.

  3. 03

    Dedicated representations for speech and 3D

    Speech synthesis first converts text into intermediate acoustic features with an acoustic model, then restores the waveform with a vocoder; 3D uses neural radiance fields (NeRF) or 3D Gaussian splatting to represent a scene, with diffusion models generating or editing those representations.

  4. 04

    Toward a unified model

    The trend shifts from one model per modality to a single backbone for all. The cost is that modalities differ sharply in compute demand, sampling rate and evaluation, so unified training requires careful balancing of data mixes and loss weights.

Abb. 3

Video is split into a grid of frames × spatial patches; each column forms a tubelet spanning time, the basic unit for modelling motion in text-to-video

0.20.70.30.10.240.680.330.120.270.650.360.140.30.630.40.17

Anwendungsfelder

  • Text-to-video, video editing, and animation or VFX assistance
  • Speech synthesis, voice cloning and music generation
  • 3D content creation: reconstructing and generating objects, scenes and digital humans
  • Unified assistants: one model for both understanding and generation, enabling cross-modal dialogue

Häufige Missverständnisse

  • The hard part of video generation is not per-frame quality but temporal consistency: flicker, sudden object changes and physically implausible motion are the main failure modes.
  • Evaluating 3D generation is far harder than 2D: 2D metrics cannot measure geometric consistency, and ground truth is often hard to obtain.
  • Unifying modalities does not simply add capabilities. Merging backbones can sacrifice single-modality peak performance, and negative transfer between modalities is real.

Schlüsselbegriffe

Spatiotemporal patch
A local unit spanning frames in video, used to model motion
Vocoder
The component that turns acoustic features back into a waveform
NeRF
A neural network representing a scene’s radiance field for novel-view synthesis
3D Gaussian Splatting
Representing a scene with many 3D Gaussian ellipsoids for fast rendering

Weiterführende Literatur