Imagem para vídeo
Colocar uma imagem estática em movimento
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes an image, optionally with a motion description, and outputs a clip that starts from it. The first frame is usually tightly constrained to the input, and later frames extrapolate motion and camera movement from there. Unlike text-to-video it has a definite visual starting point, which raises the bar for preserving subject identity and appearance.
Como é feita tecnicamente
A common approach encodes the input as a condition, either pinning the first frame of the diffusion process or injecting it as a reference, then generates the following frames; another route extends keyframes with an image model and interpolates in between. Motion magnitude and camera control are set by explicit strength parameters or trajectory conditions, while physical plausibility is learned implicitly from motion priors in the data.
Produtos representativos
8Stable Video Diffusion
2023Transforma uma imagem estática em um vídeo curto com um modelo de difusão
Kling
2024Um modelo de vídeo curto para texto-vídeo e imagem-vídeo
Hailuo
2024Um modelo de vídeo curto focado em seguir instruções e na linguagem de câmera
Dream Machine
2024Gera vídeos curtos com movimento a partir de texto ou de uma imagem
Veo
2024Gera clipes de vídeo em 1080p com planos coerentes
Gen-3 Alpha
2024Um modelo texto para vídeo altamente controlável para cinema e publicidade
Seedance
2024Um modelo de geração de vídeo voltado à narrativa em vários planos
Sora
2024Gera vídeo coerente de até cerca de um minuto a partir de uma descrição
Organizações relacionadas
Usos típicos
- Animating photos and memory clips
- Product showcases and animated ad visuals
- Motion pre-vis from key art and storyboards
- Shot drafts for games and film
Como avaliar se funciona bem
- FVD
- Distribution gap between generated and real video
- First-frame fidelity
- Agreement between the first frame and the input
- Human preference
- Pairwise judgement of motion naturalness and subject retention
Limites e dificuldades
- Identity drifts after the first frame, with faces and clothing changing first
- Large camera moves break background geometry: straight lines bend and buildings misalign
- Occlusion and parallax are often wrong, inverting front-back relations
Conceitos por trás
Modelos de difusão
Aprenda mil pequenos passos de remoção de ruído e criará uma imagem a partir de ruído puro
Difusão latente e controle condicional
Fazer difusão não sobre pixels, mas dentro de um espaço semântico comprimido
Geração multimodal
Um único modelo que aprende a falar, desenhar, mover-se e até modelar o mundo 3D