Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS ES IST
Veo is a text-to-video model Google DeepMind announced in May 2024. It accepts text or image input, generates 1080p video longer than a minute, and supports several cinematic styles and camera controls. Veo is also used in product settings such as YouTube’s short-form tools. Its weights are not open.
Warum es wichtig ist
It put high resolution, longer duration and controllable camera work into a single video model and fed it directly into Google’s product line, making it a major entry in the 2024 text-to-video race.
Wichtige Eckdaten
- Resolution
- 1080p
- Duration
- Over 60 seconds (first disclosure)
- Input
- Text, image
- Open weights
- No
Fähigkeiten
Verwandte Konzepte
Diffusionsmodelle
Lerne tausend kleine Entrauschungsschritte, und du erzeugst ein Bild aus reinem Rauschen
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt
Infrastruktur für Training und Inferenz
Der Speicher bestimmt, wie groß ein Modell sein darf, die Kommunikation, wie lange das Training dauert
Vergleichbare Produkte
Sora
2024Erzeugt aus einer Beschreibung kohärentes Video von bis zu etwa einer Minute
Gen-3 Alpha
2024Ein gut steuerbares Text-zu-Video-Modell für Film und Werbung
Kling
2024Ein Kurzvideo-Modell für Text-zu-Video und Bild-zu-Video
Dream Machine
2024Erzeugt kurze Videos mit Bewegung aus Text oder einem Bild
Seedance
2024Ein Videogenerierungsmodell für mehrteiliges Erzählen
Synthesia
2019Text eingeben, Video mit sprechendem Avatar erhalten