入力テキスト画像動画
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
これは何か
Veo is a text-to-video model Google DeepMind announced in May 2024. It accepts text or image input, generates 1080p video longer than a minute, and supports several cinematic styles and camera controls. Veo is also used in product settings such as YouTube’s short-form tools. Its weights are not open.
なぜ覚えておく価値があるか
It put high resolution, longer duration and controllable camera work into a single video model and fed it directly into Google’s product line, making it a major entry in the 2024 text-to-video race.
主な仕様
- Resolution
- 1080p
- Duration
- Over 60 seconds (first disclosure)
- Input
- Text, image
- Open weights
- No
対応する能力
関連する概念
同種の製品
Sora
2024OpenAI
一文の説明から最長約1分の一貫した動画を生成する
モデル クローズド
テキスト画像動画
Gen-3 Alpha
2024Runway
映像・広告向けの制御性の高いテキスト動画生成モデル
モデル クローズド
テキスト画像動画
Kling
2024Kuaishou (Kling)
テキストからも画像からも動画を生成できるショート動画モデル
モデル クローズド
テキスト画像動画
Dream Machine
2024Luma AI
テキストや画像から動きのある短い動画を生成する
モデル クローズド
テキスト画像動画
Seedance
2024ByteDance (Seed)
マルチショットの物語表現を狙った動画生成モデル
モデル クローズド
テキスト画像動画
Synthesia
2019Synthesia
テキストを入力するとアバターが話す動画ができる
アプリ クローズド
テキスト動画