Sector: de-slopping AI-driven research

Status: as an AI Safety researcher working with LLMs daily over the last 2 years, I think I found some best practices and synthesized them into a single "Sector workflow" repo to share with others. More detailed confidence levels .

Sector was created to avoid very specific failure modes that I'm sure many researchers experience with LLMs too, maybe even convincing us that LLM-driven research is always slop (in hands-off mode) or takes more out of you than it gives (hands-on mode). But I think I finally converged to a sweet spot between those two modes that may be generally useful to people, so let me walk you through it.

What Sector is

In 1 sentence: Sector is an agentic-engineering workflow that tells your LLM agent: "not just execute the next prompt, but also ensure that your work is a) easy to understand b) easy to verify c) engaging enough to keep the user in the driver's seat".

In 4 bullets, Sector is based on the following principles, from most to least obvious:

  1. Help the agent understand what you want:
    • Goals, what good looks like, preferred level and style of explanation
  2. Make the agent's work itself understandable, and easy to verify:
    • Plan out the task → agent implements → shows the result (not just reports about it) and confesses about ad hoc choices
  3. Next session remembers everything:
    • All proofs of work are traceable back to the producing code & commits; new agents get quickly on board with previous work
  4. Be pro-active: choose additional ways to understand and participate:
    • Lessons, games, hands-on contributions and visual navigation as your dopamine savers

The main principle: you only take what you need (with your setup agent helping you understand & choose what you need). Sector is multi-level (each level 1-4 corresponding to one of the above principles) and earlier layers don't depend on later ones. So it can be as lightweight as 1 additional block in your AGENTS.md (or CLAUDE.md) file, or as big as a whole eco-system of linked artifacts, that makes planning and verifying tasks much easier. Everything is template-based and customizable: each level only defines a shape and default templates, which you're free to adjust for your specific preferences.

Sector lives in a GitHub repo: the README has the implementation & setup details, and below I give a high-level overview.

Level-1: Agents don't understand what result is good

You ever launched a long-running session, maybe trying to attack a problem from different sides, only to discover 1 day later that all of your agent-obtained results are garbage? One of the main problems is "evaluation lens misalignment": agents often evaluate a result using different criteria than you would. So one of the basic things Sector does:

Lets you specify "what good results look like" globally, so that agents better understand what specifically you're looking for without you having to repeat it every time.

This may sound obvious or stupid, but this kind of simple specification (as short as 2-3 sentences) has made my long autonomous runs much more fruitful.

Level-2: Agents make stupid assumptions about your experiments

Even if you're aligned on "what good looks like" on a high-level, many problems are complex enough that tiny details can easily ruin their solutions -- and the agents are often bad at tiny details. The conventional answer is "so write all the code by yourself", but it's not the only solution: you can think through the tiny details without writing any code.

This is why planning sessions exist, and unlike native implementations in Claude Code/Codex where agents ask you only the 3-4 most high-level questions, Sector's planning is multi-round (inspired by Matt Pocock's grilling skills), and lets you dive much deeper into the experiment design. This results in complete, linear specifications, that you can also use later to review the work:

So you don't only see the results, but also a complete methodological chain leading to them, including all the ad-hoc choices the agent makes that can invalidate the results. These choice reveals are called Confessions and it's the number one thing that helped me catch methodological problems before it was too late. You can see it as colored boxes on the example below:

The upper part (with colored text) comes from the corresponding plan, and separates the snippets you explicitly confirmed during interview (green) from the ones the agent added on its own (amber), to help you focus on the new parts. The plan also includes other sections that have made my long-running sessions much more effective: experiment-specific evaluation metric, decision gates (e.g. what to do if the metric is low/high), when to stop or iterate further, and so on.

Level-3: The work and ideas leading to it are never lost

Research involves a lot of book-keeping, and while useful (e.g. to double check your results), it can often get boring enough for us to skip recording important details like what code version (commit ID) produced the current results, what CLI args etc. And it always backfires at the worst moment, e.g. being unable to locate the exact run that produced the headline results when you want to perform a targeted ablation during rebuttal. Sector does all of this routine book-keeping for you, in two places:

  1. Each report points exactly to how its results (artifacts) were produced: the config, the script, the commit ID and where the output was saved
  2. A branch card per workstream links the reports that still matter, in order, and records where things stand, what is established (with evidence), the landmines to avoid, and a Graveyard of dead ends with why they died (so a failed idea isn't retried a month later.)

The outcome is that it becomes much harder for the agent to fool you. The agent's goal becomes not only to deliver a nice-looking result, but also to provide a full chain behind it: from how it was produced exactly, to how it was judged (against the evaluation scheme locked in the plan), to where it was saved on disk. So you can always replicate a result and review the code behind it (by yourself or with another LLM).And the branch card just ensures that all of the important reports & takeaways from them are never lost.

All of this is maintained automatically by your agent, so you can recover the missing state just by prompting like

We're in workstream X, check how exactly we obtained figure Y. I want a variant with Z changed...

Building on existing work no longer requires re-explaining the same ideas again (or hoping the agent infers them from bare code):

We're in workstream X, I now want to run an experiment based on the data we obtained from run Y, using same evaluation lens as in experiment Z

And in case the workstream becomes too big (or too many parallel workstreams pile up), I provide a minimal UI layer, the Starmap, to help you orient yourself:

Level-4: Staying in the driver's seat

You can outsource your thinking, but you can't outsource your understanding.
(C) Andrej Karpathy

LLM-assisted research often goes so fast that your head just cannot keep up with all the new results coming in, even with careful planning. Combined with the constant incentive to "run even more comparisons/improve the metric further", you start missing the core research component of critically evaluating existing results (how they can be wrong, have you missed something important etc.) This is why Level-4 exists, allowing you to slow down at the right moments and engage with the results on proper level of detail.

Sector incentivizes this engagement by letting you be pro-active in a few activities, most of them game-like:

  • During review: Sabotage Hunt and Check My Summary. Instead of reading passively, hunt the errors planted in a walkthrough the agent writes for the game, or explain the result in your own words and get scored against the sources (works great alongside a research log that only you write, which you can incrementally update with new key artifacts + your summaries)
  • When the review surface becomes too big: a lesson. A linear walkthrough of the experiment code, step by step from the inputs to the numbers, grounded in the real code. This is the one I use most.
  • While planning: Queen's Move. Write the key piece yourself (a function, a notebook cell), agree its checks in the plan, and the agent builds the rest of the experiment around it. Earn score from the #checks passed, minus hints opened.

Red-teaming a headline result (inspired by Neel Nanda's advice on research). Not a game, but part of research mode by default: once you've checked the report, the agent offers to attack the result together with you, simplest reasons first. How could this be false? What alternative explanations fit the same data? What are the cheapest controls that would rule each one out? The controls you decide are often good next plan candidates.

In one picture, with the review loop from Level 2:

These are selected examples; the full menu (five games and two lessons) is in the add-ons guide.

Doesn't this add a lot of overhead?

In my experience, not that much:

  • Plan sessions start automatically on every major implementation request
    • Unless it's a minor follow-up/bug-fixing of existing implementation, then the agent just operates normally
  • Once the agent implements something, it automatically updates the branch card and walks you through the new report/proofs of work
  • Games are suggested automatically (by default, unless you disable it)

And it's also perfectly fine to skip some Sector ceremony if you're in a rush (I do it myself too). The basic way is to put SECTOR-OFF flag anywhere in your prompt, turning Sector off for this & subsequent prompts until you type SECTOR-ON again. Alternatively, if you say something like "I only have 5 minutes to plan & launch this", the agent will only ask you 1-3 most important implementation-changing questions. Reading full reports is not necessary too if you can judge the work from 1-2 main proofs of work (although I'd still recommend checking top Confession blocks), and games are optional. So in practice, you don't really lose any speed when it's important. Only when you have enough time, on the most central parts of the project, where it's most reasonable to slow down, plan carefully and understand the agent's work better -- that's where you can benefit the most from Sector.

But does this actually work?

Here's my honest summary:

  • As of the launch date, I'm the only person who's been using Sector for around 6 months now -- getting feedback from others is one of the main points of this launch (the form)! Initially I created it for my programming-heavy research projects in AI Safety and Interpretability (like this one), and later also used it for my personal productivity app
  • I can therefore vouch for how things work in 1-person projects developed from scratch over a period of 3 to 12 months, using frontier OpenAI and Claude models throughout; how it stands in bigger research teams sharing a repo is one of the yet-unknown things I'm curious to learn about! (see the FAQ for my initial thoughts)

  • Levels 2 and 3 ("plan -> review" loop with the underlying .sector ecosystem of notes) are the ones that I have most confidence in: they existed right from Sector's seed and have been tremendously useful in helping me plan & successfully execute complex experiments, also when under time pressure. It also kept my research progress (ideas, code, experiment results) well-organized and easy to build on.
  • Level-1 is a more recent addition. Specifying what good results look like, once, for every agent, did make my long autonomous runs noticeably more fruitful.
  • Level-4: least tested part. I do believe in it as one of the most future-proof ideas of Sector, but the implementation of it is probably far from ideal. I've mostly tested lessons in forms of walkthroughs and the Sabotage Hunt and Check My Summary games played on top of them to capture the obtained understanding (and I've been enjoying it); the other games (e.g. Queen's Move) are closer to idea sketches, and to-be refined with my & your future feedback.

As a result, researching with LLMs became less like a lottery for me and more like a strategic game, where everything is transparent, you always control the state and are rewarded for good decisions and attention to detail.


I suspect many of us have converged on similar tricks independently, so I'd love to hear how they compare to Sector. Especially if they come from failure modes not described here. So let's discuss and get better at this mess together :D

P.S. If you'd like to try Sector on a real project, the README has what this post skips: the two-step Quick start (when you run setup, say it's a research project and the agent uses the research templates), what daily use looks like, and an FAQ. And once you've tried it, I'd love to hear how it went: this 1-minute form.

  1. Of course, evaluating results is still your final call, but fixing the intermediate evaluation lens allows the agent to iterate on the solution within stated constraints like hyperparameter sweeps. This rules out boring failure modes, like something not working because of one misspecified parameter.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论