이미지-비디오 생성
정지 이미지 한 장을 움직이게 한다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes an image, optionally with a motion description, and outputs a clip that starts from it. The first frame is usually tightly constrained to the input, and later frames extrapolate motion and camera movement from there. Unlike text-to-video it has a definite visual starting point, which raises the bar for preserving subject identity and appearance.
기술적으로 구현하는 방법
A common approach encodes the input as a condition, either pinning the first frame of the diffusion process or injecting it as a reference, then generates the following frames; another route extends keyframes with an image model and interpolates in between. Motion magnitude and camera control are set by explicit strength parameters or trajectory conditions, while physical plausibility is learned implicitly from motion priors in the data.
대표 제품
8Stable Video Diffusion
2023확산 모델로 한 장의 정지 이미지를 짧은 영상으로 만든다
Kling
2024텍스트와 이미지 모두로 영상을 만드는 숏폼 모델
Hailuo
2024지시 준수와 카메라 워크에 집중한 숏폼 모델
Dream Machine
2024텍스트나 이미지로 움직임이 있는 짧은 영상을 생성한다
Veo
20241080p의 일관된 숏 영상을 생성한다
Gen-3 Alpha
2024영상·광고를 겨냥한 제어성 높은 텍스트-비디오 모델
Seedance
2024멀티 숏 내러티브를 겨냥한 영상 생성 모델
Sora
2024한 줄 설명으로 최대 약 1분 길이의 일관된 영상을 만든다
관련 기관
대표적 용도
- Animating photos and memory clips
- Product showcases and animated ad visuals
- Motion pre-vis from key art and storyboards
- Shot drafts for games and film
성능을 평가하는 방법
- FVD
- Distribution gap between generated and real video
- First-frame fidelity
- Agreement between the first frame and the input
- Human preference
- Pairwise judgement of motion naturalness and subject retention
경계와 난점
- Identity drifts after the first frame, with faces and clothing changing first
- Large camera moves break background geometry: straight lines bend and buildings misalign
- Occlusion and parallax are often wrong, inverting front-back relations