Text zu 3D
Aus einem Satz ein 3D-Modell machen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a text description and outputs a 3D asset: a mesh, a textured model or a renderable 3D representation. The output is neither an image nor a video but geometry that can be viewed from any angle and placed in a scene. Unlike image-to-3D it has no reference image at all, and unlike text-to-image its output carries a genuine third dimension.
Wie sie technisch umgesetzt wird
One route uses 2D generative models as supervision: a 2D diffusion model scores renderings from many viewpoints and that score optimises a 3D representation such as a neural radiance field or a Gaussian splat — score distillation. Another route trains a generator directly on large 3D datasets and emits mesh and texture in one or several steps. Output meshes often need retopology and decimation before entering standard rendering or game pipelines.
Repräsentative Produkte
2Beteiligte Organisationen
Typische Verwendungen
- Draft props and sets for games and film
- Quick 3D showcases for products
- Turning industrial concepts into physical form
- 3D illustration for education
Wie sie bewertet wird
- CLIP similarity
- Semantic agreement between rendered views and the prompt
- Chamfer distance
- Mean surface-point distance between generated and reference meshes
- Human rating
- Ratings for shape completeness and texture quality
Grenzen und schwierige Punkte
- Generated meshes often come out as fragmented triangles and need heavy retopology before use
- Backfaces and interiors are guessed, producing hollows and intersections when you orbit the model
- Material, lighting and geometry are not properly separated, so the look breaks when exported to another engine
Konzepte dahinter
Diffusionsmodelle
Lerne tausend kleine Entrauschungsschritte, und du erzeugst ein Bild aus reinem Rauschen
Multimodale Generierung
Ein Modell, das sprechen, zeichnen, sich bewegen — und sogar die 3D-Welt modellieren lernt
Latente Diffusion und konditionale Steuerung
Diffusion nicht über Pixel, sondern in einem komprimierten semantischen Raum