For robotics

Real-time visual and spatial understanding for robotics.

Bring real-time VLM inference to robots that need to understand their surroundings and interact with objects. Connect language instructions with scene understanding and affordance reasoning to inform object selection, grasp planning, placement, and task verification.

Illustrative applications
  • Tool handle

    Identify the part used to hold and operate the tool.

  • Package surfaces

    Consider opposing surfaces as candidate grasp regions.

  • Placement area

    Relate the object to its intended destination.

Highlighted regions show interaction context.

From perception to interaction

Understand the scene in the context of the task.

01

Object grounding

02

Spatial reasoning

03

Affordance reasoning

04

Task verification

From what the robot sees to what it does.

Generated robotics scene: a two-finger gripper lowers onto a claw hammer lying on a workbench, closes around the wooden handle, and lifts it.
Grasping a hammer by its handle
Generated robotics scene: seen from a humanoid robot's head camera, its right hand sets a small white box inside an area marked with yellow tape on a workbench.
Placing a box inside a marked area
Wrist-camera recording: a robot gripper picks up a pill bottle from a glass table, carries it, and drops it into a paper cup.
Dropping a bottle into a cup, from the wrist camera