Image-to-3D & 3D Reconstruction
Recover 3D structure from one or several photos
WHAT THIS CAPABILITY MEANS
Takes one or several images and outputs a 3D representation — mesh, point cloud, depth map or renderable radiance field. With multiple views the geometry is constrained jointly by the parallax across photos and is more trustworthy; with a single image, unobserved parts can only be inferred from priors. Unlike text-to-3D both geometry and appearance are anchored by the input images.
How it is done
The multi-view route estimates camera poses, triangulates, and refines dense depth and surfaces; neural radiance fields encode the scene as a differentiable volumetric field aligned to input views by differentiable rendering, and 3D Gaussian splatting represents it as oriented translucent ellipsoids that render faster. Single-image routes rely on priors learned from large 3D datasets, or on score distillation from a 2D diffusion prior.
Representative products
3Stable Video 3D
2024Reconstructs multi-view and 3D shape of an object from a single image
Genie
2023Generates textured 3D models from text within seconds
Stable Video Diffusion
2023Turns a single still image into a short video with a diffusion model
Organizations involved
Typical uses
- Digital archiving of artefacts and buildings
- Spatial perception for robots and driving
- Turning product photos into 3D displays
- Digital twins of real locations for film and games
How it is evaluated
- Chamfer distance
- Mean point distance between reconstructed and true surfaces
- F-score
- Balance of precision and recall within a distance threshold
- Novel-view PSNR / SSIM
- How close renderings from unseen views come to real photos
Limits and hard parts
- Single-image reconstruction guesses entirely in unseen regions, and back surfaces appear invented once you rotate
- Reflective, transparent and textureless surfaces are hard to match, and reconstruction fails broadly there
- Output scale and metrics are inaccurate, so it cannot feed manufacturing or engineering measurement directly
Concepts behind it
Image Representation
To a machine, a photo is nothing but stacked grids of numbers
Diffusion Models
Learn a thousand tiny denoising steps, and you can build an image from pure noise
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world