Figure 02
Un humanoïde de deuxième génération piloté par un modèle vision-langage-action
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE C'EST
Figure 02 is the second-generation humanoid robot Figure AI released in August 2024, designed for settings such as industry. It pairs a vision-language-action model with robot hardware so the robot can understand spoken instructions and carry out grasping and carrying tasks. Figure worked with OpenAI and demonstrated talking to the robot by voice and giving it tasks.
Pourquoi il compte
By connecting a vision-language-action model and a conversational speech model to a physical robot, it offers a direct demonstration of models entering the physical world.
Caractéristiques clés
- Form
- Bipedal humanoid robot, battery-powered
- Perception
- Camera vision
- Control
- Vision-language-action model
Concepts liés
Apprentissage par renforcement profond
Laissez un réseau de neurones décider directement depuis les pixels, stabilisé par des idées anciennes
Génération multimodale
Un seul modèle qui apprend à parler, dessiner, bouger — et même à modéliser le monde 3D
Détection d’objets
De « qu’y a-t-il dans l’image » à « quoi, où et combien »
Sécurité, alignement et injection de prompts
Le modèle optimise le proxy inscrit dans la perte, jamais ce que nous voulons vraiment : l’écart entre les deux, c’est tout le problème de l’alignement