Figure 02
A second-generation humanoid driven by a vision-language-action model
WHAT IT IS
Figure 02 is the second-generation humanoid robot Figure AI released in August 2024, designed for settings such as industry. It pairs a vision-language-action model with robot hardware so the robot can understand spoken instructions and carry out grasping and carrying tasks. Figure worked with OpenAI and demonstrated talking to the robot by voice and giving it tasks.
Why it matters
By connecting a vision-language-action model and a conversational speech model to a physical robot, it offers a direct demonstration of models entering the physical world.
Key specs
- Form
- Bipedal humanoid robot, battery-powered
- Perception
- Camera vision
- Control
- Vision-language-action model
Related concepts
Deep Reinforcement Learning
Let a neural network decide straight from pixels — then hold it steady with decades-old tricks
Multimodal Generation
One model that learns to speak, to draw, to move — even to model the 3D world
Object Detection
From “what is in the image” to “what, where, and how many”
Safety, Alignment & Prompt Injection
A model optimises the proxy we wrote into the loss, never the thing we actually want — the gap between them is the whole alignment problem