Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation [R]

preview.redd.it/i2c0cdkg04uh1.png TLDR: The question we answer: how do you learn from experts with different objectives? Pooling all their data can lose their trade-offs; learning from each expert separately misses opportunities to share data. MA-BC pools demonstrations where observed actions don’t disagree, with upper and lower bounds on sample complexity. Authors: Ziyad Sheebaelhamd, Luca Viano, Volkan Cevher, Claire Vernade

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论