WHAT IT IS
Veo is a text-to-video model Google DeepMind announced in May 2024. It accepts text or image input, generates 1080p video longer than a minute, and supports several cinematic styles and camera controls. Veo is also used in product settings such as YouTube’s short-form tools. Its weights are not open.
Why it matters
It put high resolution, longer duration and controllable camera work into a single video model and fed it directly into Google’s product line, making it a major entry in the 2024 text-to-video race.
Key specs
- Resolution
- 1080p
- Duration
- Over 60 seconds (first disclosure)
- Input
- Text, image
- Open weights
- No
Capabilities
Related concepts
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Training & Inference Infrastructure
Memory decides how large a model you can train, communication how long it takes — raw compute is rarely the bottleneck
Comparable products
Sora
2024Generates coherent video up to about a minute long from a description
Gen-3 Alpha
2024A highly controllable text-to-video model for film and advertising
Kling
2024A short-video model for both text-to-video and image-to-video
Dream Machine
2024Generates short videos with motion from text or an image
Seedance
2024A video-generation model aimed at multi-shot storytelling
Synthesia
2019Type text, get a talking-avatar video