Imagen a vídeo
Poner en movimiento una imagen fija
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes an image, optionally with a motion description, and outputs a clip that starts from it. The first frame is usually tightly constrained to the input, and later frames extrapolate motion and camera movement from there. Unlike text-to-video it has a definite visual starting point, which raises the bar for preserving subject identity and appearance.
Cómo se consigue técnicamente
A common approach encodes the input as a condition, either pinning the first frame of the diffusion process or injecting it as a reference, then generates the following frames; another route extends keyframes with an image model and interpolates in between. Motion magnitude and camera control are set by explicit strength parameters or trajectory conditions, while physical plausibility is learned implicitly from motion priors in the data.
Productos representativos
8Stable Video Diffusion
2023Convierte una imagen fija en un vídeo corto con un modelo de difusión
Kling
2024Un modelo de vídeo corto para texto a vídeo y imagen a vídeo
Hailuo
2024Un modelo de vídeo corto centrado en seguir instrucciones y el lenguaje de cámara
Dream Machine
2024Genera vídeos cortos con movimiento a partir de texto o una imagen
Veo
2024Genera clips de vídeo en 1080p con planos coherentes
Gen-3 Alpha
2024Un modelo texto a vídeo muy controlable para cine y publicidad
Seedance
2024Un modelo de generación de vídeo orientado a narrativas de varios planos
Sora
2024Genera vídeo coherente de hasta un minuto a partir de una descripción
Organizaciones relacionadas
Usos típicos
- Animating photos and memory clips
- Product showcases and animated ad visuals
- Motion pre-vis from key art and storyboards
- Shot drafts for games and film
Cómo se evalúa
- FVD
- Distribution gap between generated and real video
- First-frame fidelity
- Agreement between the first frame and the input
- Human preference
- Pairwise judgement of motion naturalness and subject retention
Límites y dificultades
- Identity drifts after the first frame, with faces and clothing changing first
- Large camera moves break background geometry: straight lines bend and buildings misalign
- Occlusion and parallax are often wrong, inverting front-back relations
Conceptos detrás
Modelos de difusión
Aprende mil pequeños pasos de eliminación de ruido y podrás crear una imagen desde ruido puro
Difusión latente y control condicional
Hacer difusión no sobre píxeles, sino dentro de un espacio semántico comprimido
Generación multimodal
Un mismo modelo que aprende a hablar, dibujar, moverse e incluso modelar el mundo 3D