Перейти к содержимому
Атлас ИИ

Текст в видео

Превратить фразу в короткий видеоклип

ВидеоНачальный #23
входТекстВидео

Полный текст статьи представлен на английском; заголовок и аннотация локализованы.

ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ

Takes a text description and outputs a video with a time dimension, where frames must stay coherent with one another. It adds a hard constraint over text-to-image: temporal consistency, so objects cannot deform or teleport between neighbouring frames. Unlike image-to-video there is no starting frame; everything is determined by the prompt.

Как это устроено

The mainstream extends diffusion from two dimensions to three: a diffusion Transformer models space–time patches jointly so attention across frames constrains motion; another line uses a latent video autoencoder with spatio-temporal attention. Variable length and resolution are managed by generating and stitching in chunks. Synchronising audio and controlling camera language are recent extensions.

Примеры продуктов

8

Sora

2024
OpenAI

Создаёт связное видео длиной около минуты по текстовому описанию

Модель Закрытый
ТекстИзображениеВидео

Veo

2024
Google DeepMind

Создаёт видеоклипы 1080p со связными планами

Модель Закрытый
ТекстИзображениеВидео

Gen-3 Alpha

2024
Runway

Хорошо управляемая модель текст-видео для кино и рекламы

Модель Закрытый
ТекстИзображениеВидео

Kling

2024
Kuaishou (Kling)

Модель коротких видео как из текста, так и из изображения

Модель Закрытый
ТекстИзображениеВидео

Hailuo

2024
MiniMax

Модель коротких видео с упором на следование инструкциям и язык камеры

Модель Закрытый
ТекстИзображениеВидео

Seedance

2024
ByteDance (Seed)

Модель генерации видео для многоракурсного повествования

Модель Закрытый
ТекстИзображениеВидео

Dream Machine

2024
Luma AI

Создаёт короткие видео с движением из текста или изображения

Модель Закрытый
ТекстИзображениеВидео

Synthesia

2019
Synthesia

Введите текст — получите видео с говорящим цифровым аватаром

Приложение Закрытый
ТекстВидео

Связанные организации

Типичное применение

  • Fast spots for ads and short films
  • Storyboards and animated pre-visualisation
  • Short-form social media assets
  • Shot drafts for games and virtual production

Как её оценивают

FVD
Distance between generated and real video distributions; lower is better
Motion consistency
Whether position and shape stay continuous across frames
Human preference
Pairwise comparison of visual quality and prompt adherence

Границы и трудности

  • Over long windows physics and causality break: objects vanish, merge or change identity
  • Complex interactions and hand motion distort most, especially where several objects touch
  • Clip length and resolution are compute-bound; beyond tens of seconds stitching shows seams

Концепции в основе