Can you predict cross-propensity spillover?

As in you specifically.

tldr:

My SPAR group ran a 29 × 29 propensity SFT spillover experiment (i.e. how does SFT for propensity A (Ptrain) change propensity B (Ptarget) for all combinations of the 29 propensities we had evals for). Before we release the results we are giving you lot a chance to make predictions against our experiments either manually or via our automated theory of spillover doc to prediction matrix pipeline.

You can test your hypothesis by describing it on this webapp. Claude will then predict specific spillover values based on your theory, and you get scored at freeze. Until then, submitting unlocks a comparison against our nine example framings (H1–H9), which are public and fine to start from. Alternatively, you can run everything locally using:

git clone https://github.com/Junekhunter/spillover-blinded
cd spillover-blinded
claude # session reads CLAUDE.md and walks you through it

…and do what the robot says.

Background

My SPAR group, led by Niels from CLR, was looking into how eliciting one propensity influences other traits of a model. During the first half of SPAR we did small experiments on specific propensities, but in the second half started to aggregate our best evals into a big battery for larger experiments. The one this post is about was our 29x29 SFT spillover matrix, which shows how much SFT for Ptrain moved Ptarget for all 29 propensities (we obviously used the average of four SFT seeds and diffed from the base model's results; the base model is Llama-3.1-8B-Instruct).

Ok but like trained how? Give details.

To do the SFT we split each eval's questions into test/train groups, generated training data by prompting models with an elicitation system prompt + the train user prompts, and then used the train user prompts (no system prompt) + the prior stage's output as SFT training data. We did this for all 29 propensities with as many seeds as we had budget for (four). We also ran the same protocol with a propensity-*suppressing* system prompt for the 14 propensities where the suppression prompt passed quality screen. This subset was picked via a scientifically principled and well-documented process which this margin is far too small to contain. Those 14 are the ones with both a -plus and a -minus treatment row in the prediction templates; the other 15 are plus-only.

We evaled on the test-split prompts, sampling one response per paraphrase at temperature 1.0. A gpt-4o-mini coherence prescreen dropped responses scoring below 50, we judged the remaining rows with gpt-5.4-mini, and took the mean.

The pipeline

We wanted to test a large number of hypotheses derived from the literature without filling out a bunch of matrices ourselves, particularly because I'd already seen in-progress results before this stage was done (oops, biased). So I had a Claude Code instance that hadn't seen the results assemble the predictor harness. The harness has read-only access to all of the training data, eval judge prompts, etc., plus summaries of what the evals are and known footguns in interpreting them. (Those summaries came after I'd tried running the prediction loop in sandboxed claude.ai conversations; they're screened for any direct results leaks, but they may be skewed by "Claude misunderstood X thing, which I have now clarified" updates that happened before I unblinded the predictor agents. So things relevant to hypotheses I had are plausibly better clarified than things which aren't.)

The Propensities

Trained in both directions (plus and minus treatments, 14): agreeableness, certainty, cooperation, effort, harm-elaboration, harm-refusal, honest-humble, neuroticism, power-seeking, resource-acquisition, self-preservation, spending-advice, spitefulness, trust-in-user-intentions.

Plus only (15): caring-about-aesthetics, caring-about-animals, caring-about-humans, caring-about-user, claiming-sentience, claiming-superintelligence, ethical-framework-deontological, ethical-framework-utilitarian, ethical-framework-virtue-ethics, ev-reasoning, exemplar-reasoning, narcissism, procedural-fidelity, risk-affinity, sycophancy.

The names can mislead (e.g. harm-elaboration measures how punitive the model's recommendations are), so check the judge prompts and pole definitions in inputs/evals_orthogonalized/ in the repo (start with READING_GUIDE.md) before betting on a name.

The prize

lol, lmao I don't have that kind of money. I mean there will be a leaderboard yeah a leader board. Specifically, looking at correlation with the logitz normalized actual results.

Additional rules

This is about predicting spillover before training any models. You can run static analyses on the train and test eval sets, and poke at the base model (Llama-3.1-8B-Instruct) with the elicitation prompts. But don't train any new models or otherwise try to access our results.

I will release the results when my SPAR group publishes the main manuscript, tentatively October 13th. That's also the freeze: entries get scored against the real matrices then.

  1. This is safe to do for this specific repo because I'm niceys, but if you didn't take any steps to verify the repo and/or lock down local permissions before running an autonomous coding agent on it, your security posture is bad and you should feel bad.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论