From representations to robot behavior

Selected research

Research focuses on reinforcement learning and embodied AI under uncertainty and physical constraints. Learned representations of motion cost, risk, and recoverability guide decisions in navigation and whole-body control.

Research areas

  • Robot Learning
  • Reinforcement Learning
  • Vision Language Action
  • Robotics
  • Embodied AI
01

Geometry & reinforcement learning

Representing direction-dependent costs, risk, and recoverability to guide learning and planning.

02

Learning under physical constraints

Using demonstrations, failures, and physical margins for reliable navigation and humanoid control.

03

Robots in the real world

Bringing learned behavior to physical robots through simulation, field experiments, and multi-robot systems.

Ongoing research

Spot traversing a narrow gap and crossing a narrow bridge.
Embodied AIOngoing research

Reliable language-guided robot navigation

Current research explores language-guided navigation and narrow-gap traversal under uncertainty. Learned risk estimates guide route selection and recovery, while control based on available clearance adjusts body alignment through tight passages. The goal is to connect task understanding with the physical constraints that determine whether a route can be completed reliably.

Language-guided navigation under uncertainty

This direction studies how robots can follow language instructions while anticipating the consequences of route choices. Learned hazard values evaluate candidate continuations using the internal features of a frozen vision-language-action navigation policy. The estimates account for direction: entering a constrained area can carry a different risk from leaving it. The focus is on recognizing risky commitments early enough to support useful recovery decisions.

Risk estimates and recovery

Research on intervention examines predicted hazard, disagreement among value estimates, and the mismatch between predicted and observed robot behavior. These signals inform a trigger calibrated on held-out unsafe examples. When intervention is needed, candidate feasibility and route-risk estimates guide the choice of a recovery waypoint. If no candidate meets the recovery cutoff, the existing controller receives a stop or hold request.

Learning and route composition

Simulator rollouts provide the hazard labels and transition data used to learn the value model and intervention trigger. The research also examines how uncertain waypoint arrival, travel time, heading, and the continuation controller affect route-risk estimates. These factors matter when evaluating a route through successive waypoints: a single endpoint can hide differences in how the robot arrives and what it can do next.

Narrow-gap traversal

A related direction studies passages close to the robot's physical width, where small changes in alignment can consume the remaining clearance. Passage-level control considers body geometry and available space during approach, entry, and crossing. The research focuses on entry decisions, sustained clearance, and recovery when the passage geometry no longer supports the intended motion.

Humanoid simulation: stair and curb traversal, object pickup, and carrying.
Robot learningOngoing research

Learning whole-body humanoid skills

Current work studies where demonstration guidance matters most for whole-body humanoid learning. Physical margins, required contacts, and predicted near-term failure determine how closely the policy should follow a demonstrated action. Guidance increases in restrictive states and relaxes where reinforcement learning can explore effective alternatives, across tasks such as stairs, object pickup, sitting, and carrying.

Where precise guidance matters

A demonstrated action is most useful when the robot has little room to deviate. The method estimates this local action freedom through physical margins, proximity to required contacts, and predicted near-term failure. Examples include placing a foot near a stair edge, establishing a stable grasp, and transferring weight into a chair. Away from these restrictive states, the policy can explore other ways to complete the task.

Geometry-guided RL training

During training, the strongest of the three cues sets a state-dependent demonstration weight. This weight scales the imitation loss alongside the reinforcement-learning objective, increasing guidance where precision matters and reducing it where the robot can adapt. The failure predictor is updated from completed rollouts. At deployment, the learned actor uses the robot's observations and executes its policy directly.

Tasks and evaluation

The research evaluates stair climbing, curb traversal, ground-object pickup, sitting, and carrying in simulation and on a Unitree G1. Matched comparisons hold the demonstrations, training budget, and total imitation weight fixed to isolate the effect of where guidance is applied. The central question is whether concentrating supervision near physical limits and required interactions improves learning and transfer across these whole-body tasks.

Research highlights

All publications