Texte vers vidéo
Transformer une phrase en clip vidéo
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.
Comment c'est fait
The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.
Produits représentatifs
8Sora
2024Génère une vidéo cohérente d’environ une minute à partir d’une description
Veo
2024Génère des clips vidéo 1080p aux plans cohérents
Gen-3 Alpha
2024Un modèle texte-vidéo très contrôlable pour le cinéma et la publicité
Kling
2024Un modèle de vidéo courte pour texte-vidéo et image-vidéo
Hailuo
2024Un modèle de vidéo courte axé sur le suivi d’instructions et le langage caméra
Seedance
2024Un modèle de génération vidéo visant le récit multi-plans
Dream Machine
2024Génère de courtes vidéos animées à partir de texte ou d’une image
Synthesia
2019Saisissez un texte, obtenez une vidéo avec un avatar qui parle
Organisations concernées
Usages typiques
- Fast spots for ads and short films
- Storyboards and animated pre-visualisation
- Short-form social media assets
- Shot drafts for games and virtual production
Comment on l'évalue
- FVD
- Distance between generated and real video distributions; lower is better
- Motion consistency
- Whether position and shape stay continuous across frames
- Human preference
- Pairwise comparison of visual quality and prompt adherence
Limites et points difficiles
- Over long windows physics and causality break: objects vanish, merge or change identity
- Complex interactions and hand motion distort most, especially where several objects touch
- Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams
Concepts sous-jacents
Modèles de diffusion
Apprenez mille petits pas de débruitage et vous créerez une image à partir de bruit pur
Génération multimodale
Un seul modèle qui apprend à parler, dessiner, bouger — et même à modéliser le monde 3D
Diffusion latente et contrôle conditionnel
Diffuser non pas sur les pixels, mais dans un espace sémantique comprimé