I built an open source framework for building RL environments. Named "Seahaven" after the fake town in The Truman Show.

I've been optimizing long-running agents: rewriting their prompts, tools, skills and subagents, and keeping the changes that score better. I've been working on some version of this problem for over a decade (at Apple, my own startup, now Kiln). The hard part is the eval environment. It needs realistic data, stateful writes, and a respawn from the same starting point for every single run. So I built the Seahaven framework. Production and staging don't work: they're shared, and you can't reset them. Hand-written mocks reset fine, but they don't hold state, and they aren't realistic enough to fool an agent. And for long tasks, you want to grade what the agent actually changed in the world, not read a 50-turn transcript. I built a few one-off environments. Doable, but hard, and every one rebuilt the same layer: a database per run, frozen starting states, parallel instances, clock control, a log of every change. So I pulled that layer into an open source framework. You write just the logic specific to your world. Seahaven handles the rest. Why: evals and RL. Both need the same thing: thousands of isolated agent runs, each from a known starting state, graded on what the agent changed. I've mostly used Seahaven for harness optimization with evals. I'm starting to tinker with RL, and I'd love to hear from anyone who tries it with GRPO. Example World: a fake Stripe. Stripe World has 24 tables and 155 API operations, behind the same tools as Stripe's own MCP server. It also serves Stripe's REST API, so well that the official Stripe SDK works against it unchanged. What Seahaven handles: A private world per run: each connection gets its own instance, and each instance gets its own SQLite DB, copied from a fixture in milliseconds. The agent can break anything. Fixtures: freeze starting states like small_startup or big_co , and reuse them across every run. Parallel: hundreds of instances per process. State diffs: every row the agent changed is logged, so you grade the result, not just the trace. Reproducible: same fixture, same clock, same random seed, same run. Composable: your world can include Stripe World (or any other world) to add its tools and APIs. Optimized for agents: includes the docs, linter and tests your coding agent needs to build a world. The loop can be as simple as this: for rollout in range(100): with world.instance("big_co", seed=rollout) as inst: run_agent(inst) # your agent, your harness reward = grade(inst.state()) # every row the agent changed OpenEnv Compatible + MCP + Web Console: Every world is an OpenEnv environment, so it works with Kiln auto-optimize , TRL's OpenEnv support or any other OpenEnv-compatible tool. You can publish worlds to Hugging Face. seahaven mcp serves a world to any MCP client, so you can point a local model at it. seahaven serve has a web console: open instances, call tools, and inspect state in your browser. Everything runs locally: Python 3.14+ and SQLite, no external services. I built it at Kiln, and Kiln uses it to evaluate and optimize agent harnesses. But Seahaven is standalone -- you don't need Kiln to use it. Seahaven is open source (MIT). Links Seahaven on GitHub Stripe World Docs Kiln for optimizing harnesses Which world should I build next? Happy to answer any questions. Side note: I made the video with videowright , another open source project of mine.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论