Figure 02
Um humanoide de segunda geração guiado por um modelo de visão-linguagem-ação
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE É
Figure 02 is the second-generation humanoid robot Figure AI released in August 2024, designed for settings such as industry. It pairs a vision-language-action model with robot hardware so the robot can understand spoken instructions and carry out grasping and carrying tasks. Figure worked with OpenAI and demonstrated talking to the robot by voice and giving it tasks.
Por que vale a pena lembrar
By connecting a vision-language-action model and a conversational speech model to a physical robot, it offers a direct demonstration of models entering the physical world.
Especificações-chave
- Form
- Bipedal humanoid robot, battery-powered
- Perception
- Camera vision
- Control
- Vision-language-action model
Conceitos relacionados
Aprendizado por reforço profundo
Deixe uma rede neural decidir direto dos pixels, estabilizada com ideias antigas
Geração multimodal
Um único modelo que aprende a falar, desenhar, mover-se e até modelar o mundo 3D
Detecção de objetos
De “o que há na imagem” para “o quê, onde e quantos”
Segurança, alinhamento e injeção de prompt
O modelo otimiza o proxy escrito na perda, nunca o que realmente queremos — essa lacuna é todo o problema de alinhamento