And here is the new PhD thesis of Gautham Vasan (@Gautham529), whose committee I was recently on. I thought this thesis was particularly well done, including the AVG algorithm in Chapter 6, which is a new streaming RL actor-critic (different from the Stream-X algorithms).

Title: Robots That Learn on the Fly Through Real-World Interaction
URL: gauthamvasan.com/papers/Vasan_Gautham_202609_PhD.pdf

Abstract:
A robot that learns from its own runtime experience can improve its behavior over its operational lifetime and adapt to conditions that were never anticipated. The prevailing practice in robot learning provides no such ability: practitioners train a control policy before deployment, from human-provided data or from simulated interaction, and hold it fixed afterward. When performance degrades after deployment, a human engineer may need to diagnose the failure, collect more data or refine the simulator, and retrain and redeploy the policy. Reinforcement learning (RL) from ongoing interaction could reduce reliance on humans for data collection, retraining, and redeployment.
Real-time reinforcement learning names the setting in which an agent senses, acts, and learns from ongoing interaction as wall-clock time advances for both the agent and the environment. Real-time RL on physical robots imposes three practical demands: (a) learning must operate within the robot’s limited computational and memory resources; (b) learning must reach competent performance reliably within a practical wall-clock budget; (c) learning must use rewards that are computable from the robot’s own sensing and simple to specify across tasks. I present three complementary contributions that address the three demands.
The first contribution, the Remote-Local Distributed (ReLoD) framework, supplements the robot’s limited computational resources with a remote workstation. The framework distributes action computation and learning updates between the two computers. Experiments on vision-based tasks showed that the performance of an agent using Soft Actor-Critic, a widely used batch RL algorithm that learns from stored transitions, degraded under limited onboard computation. A robot using ReLoD recovered performance by drawing on remote computation. Across repeated runs, competent policies were learned from scratch within a few hours.
The second contribution is an empirical study of the minimum-time task specification, which assigns a reward of −1 at each timestep and terminates the episode when the goal is reached. Given matched interaction budgets, final policies learned under the minimum-time specification matched or outperformed those learned under guiding reward specifications, which reward progress toward the goal. Using the minimum-time specification, four physical robots learned goal-reaching policies from raw-pixel observations within hours and without external instrumentation.
The third contribution is Action Value Gradient (AVG), a computationally cheap streaming RL algorithm that learns with neural networks. An agent using AVG updates its parameters once from each transition (an observation, action, reward, and next observation) and stores no transitions for reuse. On simulated tasks, agents using AVG outperformed agents using existing streaming RL methods, often achieving final performance comparable to batch RL methods. On physical-robot tasks, agents using AVG learned competent policies within hours on an edge device.
The three contributions address the three demands in complementary ways. ReLoD and AVG meet the computational resource demand from opposite directions: ReLoD supplements a robot’s onboard computer with remote computation when feasible, and AVG is computationally cheap, enabling on-device learning. The minimum-time specification reduces specifying a goal-reaching task to one requirement, detecting the goal from onboard sensing. Across the contributions, physical robots learned competent behavior from scratch within hours. Together, the contributions advance real-time reinforcement learning on physical robots toward the longer-term goal of continual adaptation and long-term autonomy.

Gautham is now a research scientist at Zyphra.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论