Sora
Erzeugt aus einer Beschreibung kohärentes Video von bis zu etwa einer Minute
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
Sora is a text-to-video model OpenAI first unveiled in February 2024. It tokenises video into spacetime patches and performs diffusion-style generation over that unified representation, which lets it accept different resolutions and durations. It accepts both text and image input for text-to-video and image-to-video. In its first demonstrations, clips ran up to about 60 seconds at 1080p, making it one of the few models to show minute-scale coherent video at the time.
Warum es wichtig ist
It brought minute-scale, shot-coherent video generation into public view for the first time, moving video generation beyond short clips toward consistency over longer stretches.
Wichtige Eckdaten
- Resolution
- Up to 1080p (first disclosure)
- Duration
- Up to about 60 seconds (first disclosure)
- Input
- Text, image
- Open weights
- No
Fähigkeiten
Verwandte Konzepte
Diffusionsmodelle
Lerne tausend kleine Entrauschungsschritte, und du erzeugst ein Bild aus reinem Rauschen
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt
Infrastruktur für Training und Inferenz
Der Speicher bestimmt, wie groß ein Modell sein darf, die Kommunikation, wie lange das Training dauert
Vergleichbare Produkte
Veo
2024Erzeugt 1080p-Videoclips mit kohärenten Einstellungen
Gen-3 Alpha
2024Ein gut steuerbares Text-zu-Video-Modell für Film und Werbung
Kling
2024Ein Kurzvideo-Modell für Text-zu-Video und Bild-zu-Video
Hailuo
2024Ein Kurzvideo-Modell mit Fokus auf Instruktionsbefolgung und Kamerasprache
Stable Video Diffusion
2023Macht aus einem Standbild mit einem Diffusionsmodell ein kurzes Video
Dream Machine
2024Erzeugt kurze Videos mit Bewegung aus Text oder einem Bild
Seedance
2024Ein Videogenerierungsmodell für mehrteiliges Erzählen
Synthesia
2019Text eingeben, Video mit sprechendem Avatar erhalten