WHAT IT IS
Sora is a text-to-video model OpenAI first unveiled in February 2024. It tokenises video into spacetime patches and performs diffusion-style generation over that unified representation, which lets it accept different resolutions and durations. It accepts both text and image input for text-to-video and image-to-video. In its first demonstrations, clips ran up to about 60 seconds at 1080p, making it one of the few models to show minute-scale coherent video at the time.
Why it matters
It brought minute-scale, shot-coherent video generation into public view for the first time, moving video generation beyond short clips toward consistency over longer stretches.
Key specs
- Resolution
- Up to 1080p (first disclosure)
- Duration
- Up to about 60 seconds (first disclosure)
- Input
- Text, image
- Open weights
- No
Capabilities
Related concepts
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Training & Inference Infrastructure
Memory decides how large a model you can train, communication how long it takes — raw compute is rarely the bottleneck
Comparable products
Veo
2024Generates 1080p video clips with coherent shots
Gen-3 Alpha
2024A highly controllable text-to-video model for film and advertising
Kling
2024A short-video model for both text-to-video and image-to-video
Hailuo
2024A short-video model focused on instruction following and camera language
Stable Video Diffusion
2023Turns a single still image into a short video with a diffusion model
Dream Machine
2024Generates short videos with motion from text or an image
Seedance
2024A video-generation model aimed at multi-shot storytelling
Synthesia
2019Type text, get a talking-avatar video