Should we be worried about how good AI is getting at coding autonomous drones?

Hey everyone!

We’re introducing Drone-Bench, a benchmark where AI agents code drones to complete a simple autonomous surveillance task. Drone-Bench is independent but based on Project Pilot, our work with Anthropic.

I would make a condensed version here for LW, but visuals are much better on web:andonlabs.com/evals/drone-bench

Very curious on your feedback!Discuss

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论