What it is
Caveman is a repo with two parts for coding agents. The skill makes replies terse while code, commands, paths and error messages stay verbatim. The proxy runs on your machine with your own keys and compresses what the agent reads, such as logs, CSV, YAML, JSON, test output and web pages. It works with Claude Code, Codex, Gemini CLI, Cursor, Windsurf, Cline, Copilot and 30+ more agents.
Who it's for
- Developers using coding agents who want to cut input and output token usage
- Users who only want shorter answers and can install just the skill
- Builders of their own agents who want to wrap Vercel AI SDK, LangChain, OpenAI or Anthropic calls with the same shrinking
Requirements
Requirements
- Node.js 22.13+ for the proxy and CLI
- Python 3.11+ for the Python middleware
- A supported coding agent such as Claude Code, Codex, Gemini, aider, kilo, qwen, opencode, hermes, openclaw or pi
Setup
Install the proxy (big rock)
Installs the CLI globally and runs setup, then launches your agent through caveman. After the first run, plain
claudestays caveman'd.bashnpm install -g @caveman-ai/cli && caveman setup --install caveman claude # or codex · gemini · aider · kilo · qwen · opencode · hermes · openclaw · piInstall only the skill (small rock)
Output compression alone, no proxy. Type
/caveman($cavemanin Codex) if it doesn't start on its own. Saystop cavemanto go back.bashnpx skills add JuliusBrussee/caveman -gInstall as a Claude Code plugin
Auto-starts every interactive session, subagents too.
bashclaude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman
Examples
Rank where your tokens go
bashcaveman learn # rank where your tokens go, from agent history already on disk
caveman learn implement # apply the fixes one diff at a time, only on your yesWhat it does: Analyzes agent history already on disk, then applies suggested fixes one diff at a time, only with your approval.
Compress noisy command output and web pages
bashcaveman shrink -- pnpm test # compress noisy command output
caveman browse <url> # a compressed web page instead of a 15,000-token dumpWhat it does: Shrinks test output and fetches web pages in compressed form for the agent.
A/B a real session
bashcaveman trial -- claude # A/B a real session on your own workWhat it does: Compares a real Claude session on your own work to check the effect.
Wrap your own agent
bashnpm install @caveman-ai/middleware @caveman-ai/sdk # TypeScript
pip install 'caveman-middleware[langchain]' caveman-sdk # Python 3.11+What it does: Installs the middleware for TypeScript or Python, wrapping an existing Vercel AI SDK, LangChain, OpenAI or Anthropic call.
Pros & cons
Pros
- Pro:The proxy cut whole-session input tokens by 33.2% across 54 Claude Code runs, with 18 of 18 answers right
- Pro:Code, commands, paths, error messages, negations, numbers and units stay verbatim, and your prompts are never rewritten
- Pro:Works with 30+ agents, with a skill-only install if you don't want the proxy
- Pro:The README publishes its losing cases, such as HTML saving nothing and 9.9% worse in the session run
Cons
- Con:The proxy has no HTML compressor yet, so the HTML session used 9.9% more tokens
- Con:The CLI sends usage stats by default (commands run, token counts, random install ID, OS, IP) until you turn it off
- Con:If you pay per request rather than per token (e.g. GitHub Copilot premium requests), shorter answers cost the same, so the README says to skip it
Images
