画像から動画生成
一枚の静止画を動かす
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes an image, optionally with a motion description, and outputs a clip that starts from it. The first frame is usually tightly constrained to the input, and later frames extrapolate motion and camera movement from there. Unlike text-to-video it has a definite visual starting point, which raises the bar for preserving subject identity and appearance.
技術的にどう実現するか
A common approach encodes the input as a condition, either pinning the first frame of the diffusion process or injecting it as a reference, then generates the following frames; another route extends keyframes with an image model and interpolates in between. Motion magnitude and camera control are set by explicit strength parameters or trajectory conditions, while physical plausibility is learned implicitly from motion priors in the data.
代表的な製品
8Stable Video Diffusion
2023拡散モデルで一枚の静止画を短い動画に変える
Kling
2024テキストからも画像からも動画を生成できるショート動画モデル
Hailuo
2024指示追従とカメラワークを重視したショート動画モデル
Dream Machine
2024テキストや画像から動きのある短い動画を生成する
Veo
20241080pでショットの一貫した動画クリップを生成する
Gen-3 Alpha
2024映像・広告向けの制御性の高いテキスト動画生成モデル
Seedance
2024マルチショットの物語表現を狙った動画生成モデル
Sora
2024一文の説明から最長約1分の一貫した動画を生成する
関連する組織
代表的な用途
- Animating photos and memory clips
- Product showcases and animated ad visuals
- Motion pre-vis from key art and storyboards
- Shot drafts for games and film
どう評価するか
- FVD
- Distribution gap between generated and real video
- First-frame fidelity
- Agreement between the first frame and the input
- Human preference
- Pairwise judgement of motion naturalness and subject retention
限界と難しさ
- Identity drifts after the first frame, with faces and clothing changing first
- Large camera moves break background geometry: straight lines bend and buildings misalign
- Occlusion and parallax are often wrong, inverting front-back relations