Figure 02
Un humanoide de segunda generación movido por un modelo de visión-lenguaje-acción
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ ES
Figure 02 is the second-generation humanoid robot Figure AI released in August 2024, designed for settings such as industry. It pairs a vision-language-action model with robot hardware so the robot can understand spoken instructions and carry out grasping and carrying tasks. Figure worked with OpenAI and demonstrated talking to the robot by voice and giving it tasks.
Por qué merece la pena recordarlo
By connecting a vision-language-action model and a conversational speech model to a physical robot, it offers a direct demonstration of models entering the physical world.
Especificaciones clave
- Form
- Bipedal humanoid robot, battery-powered
- Perception
- Camera vision
- Control
- Vision-language-action model
Conceptos relacionados
Aprendizaje por refuerzo profundo
Deja que una red neuronal decida directamente desde píxeles, estabilizada con ideas antiguas
Generación multimodal
Un mismo modelo que aprende a hablar, dibujar, moverse e incluso modelar el mundo 3D
Detección de objetos
De “qué hay en la imagen” a “qué, dónde y cuántos”
Seguridad, alineación e inyección de prompts
El modelo optimiza el proxy escrito en la pérdida, nunca lo que de verdad queremos: esa brecha es todo el problema de alineación