JuliusBrussee/caveman
🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman
why use many token when few do trick
Make your AI coding agent talk like a caveman. Same answers, 65% fewer output tokens. Brain still big. Mouth small.
github.com/JuliusBrussee/caveman/stargazers raw.githubusercontent.com/JuliusBrussee/caveman/main/INSTALL.md github.com/JuliusBrussee/caveman/commits/main raw.githubusercontent.com/JuliusBrussee/caveman/main/LICENSESee it · Install · Levels · What you get · Benchmarks · Ecosystem · Caveman 2
Caveman is a skill/plugin for Claude Code, Codex, Gemini, Cursor, Windsurf, Cline, Copilot, and 30+ other agents. Install once. Agent drops the filler and answers in tight caveman-speak, keeping code, commands, and errors byte-for-byte exact. You save output tokens on every reply, forever.
Before / After
Same fix. Third of the words. Nothing technical lost.
Caveman no make brain smaller. Caveman make mouth smaller. Shrinks what the agent says, not what it knows.
Install
One command. Finds every agent on your machine. Installs for each.
~30 seconds. Needs Node ≥18. Skips agents you no have. Safe to re-run.
Tip
Turn it on: type /caveman or say "talk like caveman". Turn it off: say "normal mode". On Claude Code, Codex, and Gemini it's already on from message one. No command needed.
Install for one agent, or any of 30+ others
Every agent has its own path (plugin, extension, rule file, or npx skills add). The full per-agent matrix, all flags, dry-run, and uninstall live in INSTALL.md. A few common ones:
Install broke? Open your agent in this repo and say: "Read CLAUDE.md and INSTALL.md, install caveman for me." Agent read repo, agent fix own brain. Snake eat tail.
Pick your grunt
Six levels. Switch anytime with /caveman . Level sticks until you change it or the session ends.
Note
Speak your tongue. Caveman keeps your language. Write Portuguese, caveman grunt Portuguese. Spanish, French, same. It compresses the style, never translates. wenyan mode is the exception on purpose: classical Chinese packs the most meaning per token.
What you get
Tip
On Claude Code the statusline shows [CAVEMAN] ⛏ 12.4k — that's your lifetime tokens saved, updated on every /caveman-stats. Silence it with CAVEMAN_STATUSLINE_SAVINGS=0.
Benchmarks
Real token counts from the Claude API. Average 65% output reduction across 10 prompts (range 22–87%), measured against default verbose replies. Output tokens only, committed and reproducible in benchmarks/ and evals/.
Important
Honest number warning. Caveman only shrinks output tokens. Input and reasoning tokens are untouched, and the skill itself adds ~1–1.5k input tokens per turn. So whole-session savings run smaller than the output number, and on already-terse workloads they can go net-negative. The real win is readability and speed. Cost savings are the bonus. When caveman wins, when it loses, and how to measure it yourself: docs/HONEST-NUMBERS.md.
Turns out short isn't just cheaper. A March 2026 paper, Brevity Constraints Reverse Performance Hierarchies in Language Models, tested 31 models and found that constraining large models to brief answers improved accuracy by ~26 points on some benchmarks. Sometimes less word = more correct.
caveman-compress receipts — real memory files, cutting input tokens forever
Every session after, that file loads ~46% smaller. Input tokens saved forever, not just one reply.
The whole cave
Five tools, one idea: agent do more with less.
Also: five sibling skills, one install
JuliusBrussee/skills — works in Claude Code, Cursor, Gemini, Cline, Copilot, 40+ agents:
🦞 Teach the lobster brevity — OpenClaw integration
OpenClaw is a self-host gateway: one box, many agents inside, wired to Slack / Discord / iMessage / Telegram. Lobster strong. Lobster smart. Lobster also talk a lot.
Same installer, scoped to one agent:
Two things happen, no more: a caveman skill lands in the workspace, and a tiny marker-fenced block is appended to SOUL.md (OpenClaw injects it every turn, so the lobster is terse from message one — no /caveman per session). Custom path? OPENCLAW_WORKSPACE=/your/path. Uninstall with the same line plus --uninstall; your other workspace content stays untouched. Lobster claw still sharp. Lobster mouth now small.
Caveman 2
Caveman make token small. Caveman 2 make it provable.
Today's savings numbers (including /caveman-stats) are local estimates. Caveman 2 measures and verifies them across a whole team — real receipts, real dashboard, real proof the tokens went down. Building it now.
Join the waitlist → caveman.so
How it works
1. Install drops a skill file into your agent.
2. Skill tells agent: drop filler, keep substance, use fragments — but never touch code, commands, or errors.
3. On Claude Code, a hook writes a tiny flag file each session, so the agent talks caveman from message one without /caveman.
4. /caveman-stats reads your session log, counts tokens saved, writes the number to your statusline.
5. /caveman-compress rewrites memory files (like CLAUDE.md) so every future session starts with a smaller context. Save tokens forever, not just once.
Hook architecture, file ownership, and CI sync are documented for maintainers in CLAUDE.md.
Privacy
Caveman no phone home. No telemetry, no analytics, no accounts, no backend. After install, zero network calls — the skill is a prompt, the hooks are local scripts, and /caveman-stats reads a log already on your disk. Install-time fetches (GitHub plus your agents' own registries) are spelled out in SECURITY.md.
Sponsors
Caveman free forever. Sponsors keep the rock sharp.
atlascloud.aiAtlas Cloud — full-modal AI inference platform, one API.
Want your rock here? → Sponsor caveman
Star this repo
Caveman save you token, save you money. Star cost zero. Fair trade. ⭐
Docs: Install matrix · Honest numbers · Contributing · Maintainer guide · Issues Also by Julius Brussee:Revu — local-first macOS study app with FSRS spaced repetition (revu.cards)
MIT — free like mass mammoth on open plain.