Migrating CompileBench to Harbor: standardizing AI agent evals

Standardizing AI agent evaluation with Harbor: an open-source framework for reproducible benchmarks, reinforcement learning, and collaborative evals.
评论
?
参与讨论

Standardizing AI agent evaluation with Harbor: an open-source framework for reproducible benchmarks, reinforcement learning, and collaborative evals.