Caveman

Cut 65% of your AI costs.

Your AI bill is mostly waste. We find it, cut it, and prove every dollar saved.

74kon GitHub#1 on Hacker News

≈35% kept65% cut

65% fewer output tokens · measured across 10 prompts

Developers at these companies starred Caveman.

74,000 stars on GitHub

Caveman Cloud · the platform

One platform to make your entire AI stack more efficient.

Caveman Cloud usage view — tokens, models, cache and compression share, and tokens over time, priced at public list
Usage
Caveman Cloud activity view — the same job done two ways, with the cheapest and costlier paths priced per run
Activity
real product UI · demo workspace · private beta

Caveman Cloud · optimize

Any optimization you can think of.

Model routing, auto caching, prompt compression, waste detection — every request takes the cheapest path that still does the job. Switched on at the gateway, proven in the ledger.

harness
gateway

Caveman router

Every request takes the cheapest model that still does the job.

Learn more

standard answer · ~72 tokens

The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.

caveman · ~21 tokens

New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo.
Claudeskill · cavemanRun

Caveman skill

The original — 100k+ stars on GitHub. ~10% savings on long-horizon coding tasks in JetBrains' benchmark.

Auto caching

Cache writes and reads placed automatically wherever a prompt repeats. You keep the discount.

Compression

Context compressed before it ships. The same answer comes back on fewer tokens.

Waste detection

20 detectors read your traffic and rank every dollar of waste by what fixing it returns.

Hooks into what you have

Drop-in for LiteLLM, Vercel AI SDK, LangChain — anything that speaks OpenAI.

illustrative route · savings measured per workspace, never assumed

Caveman Cloud · surfaces

Savings, wherever your team already works.

Coding agents, app SDKs, whole pipelines — point them at the gateway and the optimizations ride along. No rewrite, no new framework.

One environment variable. Every session caches its reads, compresses its context, and leaves a receipt.

Codex already speaks OpenAI — export one URL and every run is routed, cached, and receipted.

Any OpenAI-compatible agent CLI rides the same gateway. Hermes included, unchanged.

Swap the baseURL in createOpenAI. Ship the same app; pay for fewer tokens.

Point the agent's model at the gateway and every run is routed, cached, and receipted.

terminal · claudeillustrative
$export ANTHROPIC_BASE_URL=https://gw.caveman.so
$claude
Claude Code — session routed through Caveman
>fix the flaky auth test
reading tests/auth.spec.ts — cache read, 41k tokens skipped
edit applied · tests pass ✓
session receipt: cached reads billed at the cache rate
tokens saved / run · illustrative
01Verification

Numbers that survive your CFO.

Savings climb a ladder: guessed, measured, proven. A number never skips a step. Failed changes roll themselves back. Every receipt is signed.

The proof plane · inferred → replayed → verifiedsigned · auditable
A
the ladder · a number never skips a step
inferreda guess · per day

A guess, read off your own traffic. Shown per day. Never a promise.

replayedmeasured

We re-ran the fix on real requests. Measured, not guessed. Still not counted.

verifiedsigned · counted

Live on real traffic and signed into the ledger. The only number we count.

verified_savings = $0.00until active on real traffic

Every vendor shows you a big number. Ours starts at zero and only moves when the money is real.

B

Shadow first, always

Every change is measured against a baseline before it touches live traffic. If it fails, it rolls itself back.

what a gate checks
exactness

code, commands, and errors byte-for-byte

structure

JSON and tool schemas still validate

task pass rate

the work still succeeds

latency & cost

it really does spend less, and it is not slower

recordreplayshadowcanaryactive

gate ✓ at every stage → active

27 grader types · unknown graders fail closed · auto-rollback on regression
ed25519-signed, hash-chained receipts · export is manual today · automatic signing stays off until the ledger can attest complete days
06Models & research

Caveman for research.

Caveman Labs · Research division4 papers

CaveGemma

fine-tune of google/gemma-4-31B-it · QLoRA r16 · MIT · weights inherit Gemma terms

Get the weights →
27%
fewer output tokens · 193 pairs
96–100%
code-fence exactness
0.91–0.98
semantic cosine
534MB
LoRA adapter
Specimen · logs as pixelsFig. 00
Input · agent system promptOutput · −65%
07News

What we think, in full.

Read everything

Run your number

Waste scales with headcount.

The 65% is measured. The rest is your sliders, compounded.

200
$150
usage growth, month over month5%

compounded monthly across the year

projectionlist price · estimate
spend today$477.5K / yr
→ with Caveman$167.1K / yr
estimated annual savings
$310.4K
65% measured cut

$310,384 of $477,514 modeled spend

Modeled from the 65% output-token cut measured across 10 prompts, applied at list price. An estimate, not a bill.

That’s $310.4K a year on the table.

Two weeks of traffic is enough to rank every dollar of waste. Cloud is in private beta.

No spam · one email when your invite is ready