Y'all this is a sexy paper; context language models

Paper linky - Context Language Models The central idea of the paper is incredibly simple. Give a model the ability to edit its context on-the-go like a file has major benefits on task performance, context management (memory) and even computational efficiency (both wall clock and total flops). Their paper shows mostly benefits and relatively small downsides. You can try it out as a plugin for pi! In short, pros and cons Pros: Improves outcomes on long running tasks Coding and deep research tasks Open discovery problems (long horizon research tasks, /goal loops etc) Inference can become more compute-efficient and wall-clock efficient Note, this depends on a caching optimization in the inference engine Much less context bloat, meaning it's more (V)RAM efficient No more slow and unreliable compacts Cons: The cache optimization only exists for SGLang Prompt injections (including hallucinated instructions) are much less likely to be forgotten, increasing risks Requires harness customizations (authors supply a pi plugin) Some more context The approach works by modifying the harness to allow access to the context as a file. A model is allowed to edit the context as it would any other file. They've tested the approach on models as small as qwen3.6 9b, as well as on qwen3.8 27b and claude sonnet 4.6. Out-of-the-box, meaning just a small addition to the system prompt and tools to edit the context as a file, task performance, context management and efficiency measures remain approximately the same or improve by a little bit. The smaller qwen3.6 9b model in particular lost a little bit of efficiency, suggesting it works better on larger (smarter) models. Performance can be massively improved with RL training, which the authors also did. Wanna try it out? You can try it out right now if you use pi Install the plugin github.com/lolipopshock/pi-clm , this comes from the authors directly After installation, adjust settings with /clm settings : Set steering to house-brief.md (modifies the system prompt, I suppose this should be left disabled for RL'd models only, of which there are none right now) Enable "One tool per turn"; this one is important for performance Enable "Size trailer"; this one appends context usage after every tool result. Without it, models are much less inclined to modify context on-the-go for large tool calls Fin Let me know how it goes! Last, I also consulted this video by "Prompt Engineering" on YouTube in addition to the paper: youtube.com/watch

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论