Geometry & reinforcement learning
Representing direction-dependent costs, risk, and recoverability to guide learning and planning.
From representations to robot behavior
Research focuses on reinforcement learning and embodied AI under uncertainty and physical constraints. Learned representations of motion cost, risk, and recoverability guide decisions in navigation and whole-body control.
Research areas
Representing direction-dependent costs, risk, and recoverability to guide learning and planning.
Using demonstrations, failures, and physical margins for reliable navigation and humanoid control.
Bringing learned behavior to physical robots through simulation, field experiments, and multi-robot systems.
Current research explores language-guided navigation and narrow-gap traversal under uncertainty. Learned risk estimates guide route selection and recovery, while control based on available clearance adjusts body alignment through tight passages. The goal is to connect task understanding with the physical constraints that determine whether a route can be completed reliably.
This direction studies how robots can follow language instructions while anticipating the consequences of route choices. Learned hazard values evaluate candidate continuations using the internal features of a frozen vision-language-action navigation policy. The estimates account for direction: entering a constrained area can carry a different risk from leaving it. The focus is on recognizing risky commitments early enough to support useful recovery decisions.
Research on intervention examines predicted hazard, disagreement among value estimates, and the mismatch between predicted and observed robot behavior. These signals inform a trigger calibrated on held-out unsafe examples. When intervention is needed, candidate feasibility and route-risk estimates guide the choice of a recovery waypoint. If no candidate meets the recovery cutoff, the existing controller receives a stop or hold request.
Simulator rollouts provide the hazard labels and transition data used to learn the value model and intervention trigger. The research also examines how uncertain waypoint arrival, travel time, heading, and the continuation controller affect route-risk estimates. These factors matter when evaluating a route through successive waypoints: a single endpoint can hide differences in how the robot arrives and what it can do next.
A related direction studies passages close to the robot's physical width, where small changes in alignment can consume the remaining clearance. Passage-level control considers body geometry and available space during approach, entry, and crossing. The research focuses on entry decisions, sustained clearance, and recovery when the passage geometry no longer supports the intended motion.
Current work studies where demonstration guidance matters most for whole-body humanoid learning. Physical margins, required contacts, and predicted near-term failure determine how closely the policy should follow a demonstrated action. Guidance increases in restrictive states and relaxes where reinforcement learning can explore effective alternatives, across tasks such as stairs, object pickup, sitting, and carrying.
A demonstrated action is most useful when the robot has little room to deviate. The method estimates this local action freedom through physical margins, proximity to required contacts, and predicted near-term failure. Examples include placing a foot near a stair edge, establishing a stable grasp, and transferring weight into a chair. Away from these restrictive states, the policy can explore other ways to complete the task.
During training, the strongest of the three cues sets a state-dependent demonstration weight. This weight scales the imitation loss alongside the reinforcement-learning objective, increasing guidance where precision matters and reducing it where the robot can adapt. The failure predictor is updated from completed rollouts. At deployment, the learned actor uses the robot's observations and executes its policy directly.
The research evaluates stair climbing, curb traversal, ground-object pickup, sitting, and carrying in simulation and on a Unitree G1. Matched comparisons hold the demonstrations, training budget, and total imitation weight fixed to isolate the effect of where guidance is applied. The central question is whether concentrating supervision near physical limits and required interactions improves learning and transfer across these whole-body tasks.

Learning a directed navigation distance from demonstrations, failures, and interventions to guide safer decisions over long horizons.

Learning value geometry that accounts for direction-dependent motion costs and rare, high-cost failures.

Separating potential differences from path-dependent costs to structure goal-reaching under asymmetric traversal.

Using available clearance to guide a legged robot into and through passages near its physical limits.

Connecting physical and virtual robots through bandwidth-adaptive synchronization for collaborative navigation.