텍스트-비디오 생성
한 문장 설명으로 짧은 영상을 만든다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.
기술적으로 구현하는 방법
The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.
대표 제품
8Sora
2024한 줄 설명으로 최대 약 1분 길이의 일관된 영상을 만든다
Veo
20241080p의 일관된 숏 영상을 생성한다
Gen-3 Alpha
2024영상·광고를 겨냥한 제어성 높은 텍스트-비디오 모델
Kling
2024텍스트와 이미지 모두로 영상을 만드는 숏폼 모델
Hailuo
2024지시 준수와 카메라 워크에 집중한 숏폼 모델
Seedance
2024멀티 숏 내러티브를 겨냥한 영상 생성 모델
Dream Machine
2024텍스트나 이미지로 움직임이 있는 짧은 영상을 생성한다
Synthesia
2019글자를 입력하면 아바타가 말하는 영상이 나온다
관련 기관
대표적 용도
- Fast spots for ads and short films
- Storyboards and animated pre-visualisation
- Short-form social media assets
- Shot drafts for games and virtual production
성능을 평가하는 방법
- FVD
- Distance between generated and real video distributions; lower is better
- Motion consistency
- Whether position and shape stay continuous across frames
- Human preference
- Pairwise comparison of visual quality and prompt adherence
경계와 난점
- Over long windows physics and causality break: objects vanish, merge or change identity
- Complex interactions and hand motion distort most, especially where several objects touch
- Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams