मुख्य सामग्री पर जाएँ

इमेज-से-वीडियो

एक स्थिर छवि को गतिमान बनाना

वीडियोप्रारंभिक #24
इनपुटइमेजवीडियो

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes an image, optionally with a motion description, and outputs a clip that starts from it. The first frame is usually tightly constrained to the input, and later frames extrapolate motion and camera movement from there. Unlike text-to-video it has a definite visual starting point, which raises the bar for preserving subject identity and appearance.

तकनीकी रूप से कैसे

A common approach encodes the input as a condition, either pinning the first frame of the diffusion process or injecting it as a reference, then generates the following frames; another route extends keyframes with an image model and interpolates in between. Motion magnitude and camera control are set by explicit strength parameters or trajectory conditions, while physical plausibility is learned implicitly from motion priors in the data.

प्रतिनिधि उत्पाद

8

Stable Video Diffusion

2023
Stability AI

विसरण मॉडल से एक स्थिर चित्र को छोटे वीडियो में बदलता है

मॉडल खुले वेट
इमेजवीडियो

Kling

2024
Kuaishou (Kling)

टेक्स्ट-टू-वीडियो और इमेज-टू-वीडियो दोनों के लिए शॉर्ट-वीडियो मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजवीडियो

Hailuo

2024
MiniMax

निर्देश-पालन और कैमरा भाषा पर केंद्रित शॉर्ट-वीडियो मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजवीडियो

Dream Machine

2024
Luma AI

पाठ या चित्र से गति युक्त छोटे वीडियो बनाता है

मॉडल बंद स्रोत
टेक्स्टइमेजवीडियो

Veo

2024
Google DeepMind

सुसंगत शॉट के साथ 1080p वीडियो क्लिप बनाता है

मॉडल बंद स्रोत
टेक्स्टइमेजवीडियो

Gen-3 Alpha

2024
Runway

फ़िल्म और विज्ञापन के लिए अत्यधिक नियंत्रण योग्य टेक्स्ट-टू-वीडियो मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजवीडियो

Seedance

2024
ByteDance (Seed)

मल्टी-शॉट कथा-कथन के लिए वीडियो जनरेटिंग मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजवीडियो

Sora

2024
OpenAI

एक विवरण से लगभग एक मिनट तक का सुसंगत वीडियो बनाता है

मॉडल बंद स्रोत
टेक्स्टइमेजवीडियो

संबंधित संस्थान

सामान्य उपयोग

  • Animating photos and memory clips
  • Product showcases and animated ad visuals
  • Motion pre-vis from key art and storyboards
  • Shot drafts for games and film

इसका मूल्यांकन कैसे होता है

FVD
Distribution gap between generated and real video
First-frame fidelity
Agreement between the first frame and the input
Human preference
Pairwise judgement of motion naturalness and subject retention

सीमाएँ और कठिनाइयाँ

  • Identity drifts after the first frame, with faces and clothing changing first
  • Large camera moves break background geometry: straight lines bend and buildings misalign
  • Occlusion and parallax are often wrong, inverting front-back relations

इसके पीछे की अवधारणाएँ