VLA & Robot World Model

  • World-model construction for physical prediction: learns from a robot's sensor data to build a neural world model that simulates, in real time, how objects' physical properties and interactions will evolve into future states.
  • VLA-based autonomous action control: fuses natural-language commands with the world model's predicted context in an end-to-end controller that directly derives optimal robot actions in complex, dynamic environments.

Vision-Language-Action (VLA) and World-Model-based technology for autonomous physical-intelligence robot control.

▶ Watch demo
Robot arm and end-effector detection, and multi-view video generation from the world model
Object/end-effector detection and multi-view future-state generation from the world model.

3D VLP (Vision-Language-Planning)

  • Semantic 3D space construction: fuses visual information from heterogeneous multi-robot teams to build a 3D-LLM map that perceives the environment in real time.
  • Intelligent unified control: interprets natural-language commands within spatial context and automatically generates and controls optimized collaborative actions across heterogeneous multi-robot teams.

3D Vision-Language-Planning (3D VLP) technology for natural-language-command-driven autonomous heterogeneous multi-robot operation.

3D-VLP roadmap: real-time 3D-LLM map update, multi-robot embodied planning, and Sim-to-Real validation
3D-VLP roadmap — real-time 3D-LLM map update, multi-robot embodied planning, and Sim-to-Real validation.
← Back to Research