Towards smarter Vision-Language-Action models
In this document, I outline what I consider to be the main limitations of Visual-Language-Action (VLA) models and propose potential ways to address them. The architecture I suggest offers several advantages:
- Extension of VJepa: Fully self-supervised, eliminating the need to collect data through teleoperation of the robot.
- Explicit Hierarchical Representation of Sensorimotor Loops: Supports both fast and slow robot control loops, reducing latency issues common in large VLA models.
- Integration of Sensor, Motor, and Planning Capabilities: Sensors and motors are represented at multiple levels of abstraction, allowing the architecture to create a cohesive representation. This improves the robot's efficiency in determining the appropriate motor plan to interact with objects in the visual scene.




