Text zu Video
Aus einem Satz einen kurzen Clip machen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.
Wie sie technisch umgesetzt wird
The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.
Repräsentative Produkte
8Sora
2024Erzeugt aus einer Beschreibung kohärentes Video von bis zu etwa einer Minute
Veo
2024Erzeugt 1080p-Videoclips mit kohärenten Einstellungen
Gen-3 Alpha
2024Ein gut steuerbares Text-zu-Video-Modell für Film und Werbung
Kling
2024Ein Kurzvideo-Modell für Text-zu-Video und Bild-zu-Video
Hailuo
2024Ein Kurzvideo-Modell mit Fokus auf Instruktionsbefolgung und Kamerasprache
Seedance
2024Ein Videogenerierungsmodell für mehrteiliges Erzählen
Dream Machine
2024Erzeugt kurze Videos mit Bewegung aus Text oder einem Bild
Synthesia
2019Text eingeben, Video mit sprechendem Avatar erhalten
Beteiligte Organisationen
Typische Verwendungen
- Fast spots for ads and short films
- Storyboards and animated pre-visualisation
- Short-form social media assets
- Shot drafts for games and virtual production
Wie sie bewertet wird
- FVD
- Distance between generated and real video distributions; lower is better
- Motion consistency
- Whether position and shape stay continuous across frames
- Human preference
- Pairwise comparison of visual quality and prompt adherence
Grenzen und schwierige Punkte
- Over long windows physics and causality break: objects vanish, merge or change identity
- Complex interactions and hand motion distort most, especially where several objects touch
- Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams
Konzepte dahinter
Diffusionsmodelle
Lerne tausend kleine Entrauschungsschritte, und du erzeugst ein Bild aus reinem Rauschen
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt
Latente Diffusion und konditionale Steuerung
Diffusion nicht über Pixel, sondern in einem komprimierten semantischen Raum