WHAT IT IS
Kling is a video-generation model Kuaishou released in June 2024. It supports both text-to-video and image-to-video, produces clips of 5 or 10 seconds up to 1080p, and emphasises motion magnitude and visual steadiness. It was opened to users through a web product and does not disclose its weights.
Why it matters
It was an early product to offer both text-to-video and image-to-video to the public and gained wide use in 2024 for its larger motion range and steadiness.
Key specs
- Resolution
- Up to 1080p
- Duration
- 5 or 10 seconds
- Input
- Text, image
- Open weights
- No
Capabilities
Related concepts
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Training & Inference Infrastructure
Memory decides how large a model you can train, communication how long it takes — raw compute is rarely the bottleneck
Comparable products
Sora
2024Generates coherent video up to about a minute long from a description
Veo
2024Generates 1080p video clips with coherent shots
Gen-3 Alpha
2024A highly controllable text-to-video model for film and advertising
Hailuo
2024A short-video model focused on instruction following and camera language
Stable Video Diffusion
2023Turns a single still image into a short video with a diffusion model
Dream Machine
2024Generates short videos with motion from text or an image
Seedance
2024A video-generation model aimed at multi-shot storytelling
SenseAvatar
2022Generates lip-synced digital-human video from a portrait and a voice track