Изображение в 3D и реконструкция
Восстановить трёхмерную структуру по одному или нескольким снимкам
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes one or several images and outputs a 3D representation — mesh, point cloud, depth map or renderable radiance field. With multiple views the geometry is constrained jointly by the parallax across photos and is more trustworthy; with a single image, unobserved parts can only be inferred from priors. Unlike text-to-3D both geometry and appearance are anchored by the input images.
Как это устроено
The multi-view route estimates camera poses, triangulates, and refines dense depth and surfaces; neural radiance fields encode the scene as a differentiable volumetric field aligned to input views by differentiable rendering, and 3D Gaussian splatting represents it as oriented translucent ellipsoids that render faster. Single-image routes rely on priors learned from large 3D datasets, or on score distillation from a 2D diffusion prior.
Примеры продуктов
3Stable Video 3D
2024Восстанавливает множество ракурсов и трёхмерную форму объекта по одному изображению
Genie
2023За секунды создаёт текстурированные трёхмерные модели из текста
Stable Video Diffusion
2023Превращает один статичный кадр в короткое видео с помощью диффузии
Связанные организации
Типичное применение
- Digital archiving of artefacts and buildings
- Spatial perception for robots and driving
- Turning product photos into 3D displays
- Digital twins of real locations for film and games
Как её оценивают
- Chamfer distance
- Mean point distance between reconstructed and true surfaces
- F-score
- Balance of precision and recall within a distance threshold
- Novel-view PSNR / SSIM
- How close renderings from unseen views come to real photos
Границы и трудности
- Single-image reconstruction guesses entirely in unseen regions, and back surfaces appear invented once you rotate
- Reflective, transparent and textureless surfaces are hard to match, and reconstruction fails broadly there
- Output scale and metrics are inaccurate, so it cannot feed manufacturing or engineering measurement directly
Концепции в основе
Цифровое представление изображения
Для машины фотография — лишь набор наложенных сеток чисел
Диффузионные модели
Научитесь тысяче мелких шагов удаления шума — и соберёте изображение из чистого шума
Мультимодальная генерация
Одна модель учится говорить, рисовать, двигаться — и даже моделировать трёхмерный мир