For developers and engineering teams

Your coding agent.
Less expensive.That’s the job.

Your agent reads too much and repeats too much. Caveman cuts the waste around the work, with smaller context and shorter replies. Keep the coding agent you already use.

Local tools available now. No Caveman account needed.

Inside a coding sessionHow it fits
Your coding agent
CavemanOn your machine
01

Shrink bulky context

Input
02

Keep originals in reach

Recovery
03

Ask for shorter replies

Output

Your AI provider

Less text sent. Original details kept.Skill + Proxy

Claude Code · Codex · Gemini · and moreKeep your agent. Add Caveman.

01 / The problem

You pay for the work.
And everything around it.

A coding task is more than one question. Your agent reads files, checks output, and sends that growing history back to the model. The extra text adds up.

01

Too much to read

Large logs, files, and tool results travel into later turns.

02

Too much to say

Long replies cost tokens and leave more history to reread.

03

Too much to repeat

The same background work happens across steps and sessions.

02 / How Caveman helps

Less waste.
At every step.

Start with the local Proxy and Skill. Add focused browser tools, or use the SDK when building your own agents. Each method tackles a different part of the cost.

Available locally

Smaller context

Your agent sends old logs, files, and tool results back to the model again and again.

Bulky tool result
Compress + keep original
Smaller input

What Caveman does

Caveman compresses supported content before the next call. Logs, JSON, tables, code, and diffs each get a method suited to their format.

What changes for you

Less text for the model to read on later turns.

Originals stay available for exact recovery. Content passes through when compression cannot be applied safely.

Explore context compression

03 / The impact

Measured on real calls.
With the limits in view.

A published Claude Code benchmark compared direct calls with Caveman wrap + Skill on six fixed tool-output tasks, three times each.

33.2%

fewer provider-reported input tokens

18/18

exact-answer checks passed in both groups

Input tokens across three runs per task
TaskDirectWith CavemanChange
Find an outlier in a table165,82374,48455.1%
Find an error in logs148,80774,06850.2%
Check a configuration132,12471,02746.2%
Find a failing test150,377108,51427.8%
Check deployment data147,975108,93926.4%
Read a dashboard page140,687154,641+9.9%
All six tasks885,793591,67333.2%

The dashboard task used more tokens: no compression applied, but the Skill still added overhead. These are controlled tasks, not customer savings. Cache tokens are counted without price weighting, so 33.2% fewer tokens does not mean a 33.2% lower bill.

Published report, August 2026. Raw runs and the harness are not public; this result cannot yet be independently reproduced from the repository.

Read the benchmark and its limits

04 / What reliable means

The work still
has to work.

Fewer tokens only help when the task gets done. Keep the original details, test the result, and compare the complete cost.

01

Keep a way back.

The local Proxy stores original content before replacing it. Your agent can recover exact detail when it needs it.

02

Check the answer.

Run your normal tests. A smaller conversation is useful only if the fix is still correct.

03

Count the whole bill.

Include prompt overhead, cache charges, recovery calls, and retries. A smaller token count is not a guaranteed bill reduction.

Keep your agent.
Add Caveman.

Start with free local compression. Compare a real task with and without Caveman.

Token-based usage can fall. Fixed subscriptions and per-request charges may stay the same.