Can mid-training survive RL

About CaML

CaML is an alignment nonprofit working to make AI systems more compassionate toward all sentient beings with a focus on alignment midtraining (e.g. Teaching Claude Why). We build evaluation benchmarks on UK AISI's Inspect framework (TAC, MCP), run the public leaderboard at compassionbench.com, generate synthetic documents, and research which training interventions most effectively instill compassionate values. Our early results and those of others (Geodesic) show large gains from midtraining are possible. The open question here is whether any midtraining gains survive reinforcement learning and how much.


https://www.compassionml.com/careers

The project

Under what conditions do values instilled through midtraining (synthetic document fine-tuning (SDF) for values/propensities) persist after subsequent reinforcement learning, and when are they eroded or erased? We will use different claims in synthetic documents with OLMo 3’s base model and fine-tuning to test hypotheses on what drives this non-RL alignment method to sometimes be washed out and sometimes prove robust.

Deliverables include an open persistence-evaluation harness on Inspect, released model checkpoints and corpora, and at least one paper.

What you'll do

  • Help build and run the training pipeline end to end: midtraining runs on synthetic corpora, followed by RL post-training (e.g. RLVR with GRPO on verifiable tasks)
  • Implement the synthetic corpus generation pipeline, including well-matched control conditions, at the scale and quality the experiments demand
  • Run evaluations and mech-interp methods under guidance
  • Reproducible configurations, checkpointing, thorough logging, and clean experiment tracking so that every result can be traced and rerun
  • Debug issues: training instabilities, and infrastructure failures
  • Co-author papers and contribute to open-source releases

What we're looking for

  • Strong ML engineering skills: you have finetuned open-weight models at the 7B scale or larger and are comfortable with Unsloth training, Runpod and the Hugging Face ecosystem
  • A solid working understanding of RL for language models: you can implement, run, and debug an RLVR or GRPO pipeline, and you understand what the algorithms are doing well enough to notice when a run is quietly broken
  • Familiarity with the literature on RL, synthetic document midtraining and self-fulfilling (mis)alignment
  • Careful engineering habits: reproducibility, versioning, and comprehensive logging are things you do by default, not on request
  • Ability to work independently in a small distributed team and communicate clearly in writing

Nice to have:

  • A track record of public output (papers, preprints, open-source contributions, or technical posts)
  • Ideas around values training and pretraining/midtraining data persistence
  • Experience with Inspect, OLMo models, or open post-training frameworks
  • Basic interpretability skills for probing the mechanisms behind persistence and erosion
  • Familiarity with CaML’s past research

Logistics

  • Full-time for 12 months, with extension contingent on further funding
  • Remote-first, but California time zones
  • Presence in the SF Bay Area (Berkeley) is a plus but not required
  • Salary equivalent to approximately $110,000 per year as a contractor
  • Report to the co-founders and work day-to-day with our Technical Lead (Jasmine)

How to apply

Send a CV and a short note (no cover letter needed) pointing us to the single piece of your past work most relevant to this project and why you want to work with us, to hello@compassionml.com. If selected you will take a technical interview and a cultural fit interview.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论