---
title: "Caveman pricing"
description: "Free on your machine. Cloud when your team needs proof."
canonical: https://caveman.so/pricing
last-updated: 2026-10-09
---

# Caveman pricing

Free on your machine. Cloud when your team needs proof.

Caveman's agent and SDK are free and open source. Caveman Cloud sees what your agents cost, finds the cheaper way and proves it on your tasks, hosted or in your VPC.

## Free: $0 / free forever

Includes caveman /agent & /sdk. No account. Runs on your machine. For developers and the agents they build.

Free features:

### caveman /agent

For Claude Code, Codex, Gemini and 30+ coding agents.

- Output compression: The agent says less · 65% fewer output tokens across 10 prompts
- Content-aware compression: Logs, JSON, code, diffs and pages, each with its own compressor · 33% fewer input tokens
- Model routing (soon): The right model each turn, on your keys
- Waste fixes: The agent finds its worst waste and fixes it, one diff at a time
- Reusable scripts (soon): Blocks · the agent keeps the scripts it writes, checked every PR

### caveman /sdk

For LangChain, Vercel AI SDK, OpenAI and Anthropic SDKs.

- Content-aware compression: The same per-format compressors inside AI SDK, LangChain, OpenAI and Anthropic calls
- Byte-exact recovery: The model fetches the original when it needs it
- Tool-schema compression: A smaller tool catalog, same tool picks

Install the agent: `npm i -g @caveman-ai/cli && caveman setup --install`

Install the sdk: `npm i @caveman-ai/middleware`

[Install free](https://caveman.so/solutions/coding-agents)

## Team: Custom

Includes caveman /cloud. Private development. For teams that want agent costs cut, with proof.

Everything in Free, plus:

### caveman /cloud

Hosted at app.caveman.so.

- Observability: Cost per task across every agent and run · imports the logs and traces you already have
- Harness optimization (dev): Prompts, tools, caching and compression, changed automatically and shipped with proof
- Model fit per task: The cheapest model that still passes, chosen per task
- Evals: Configured by agents from your traffic, judged on your own tasks
- Agent-ready: Your coding agent drives all of it over skills and MCP

[Book a call](https://cal.com/caveman/chat)

## Enterprise: Custom / year

Includes caveman /cloud in your VPC. Annual contract. For orgs that keep data in their own cloud.

Everything in Team, plus:

### caveman /cloud

In your VPC, one Helm chart on AWS or GCP.

- Runs in your VPC: One Helm chart on your Kubernetes, AWS or GCP
- Your KMS and storage: Keys, traces and artifacts stay in your cloud account
- Signed releases: The same signed images as hosted Cloud, verifiable before you install
- Annual contract: Priced on the AI spend Caveman manages, never per seat

[Talk to us](https://cal.com/caveman/chat)

## Compare plans

| | Free | Team | Enterprise |
| --- | --- | --- | --- |
| **caveman /agent** | | | |
| Output compression | Yes | Yes | Yes |
| Content-aware compression | Yes | Yes | Yes |
| Model routing | soon | soon | soon |
| Waste fixes | Yes | Yes | Yes |
| Reusable scripts | soon | soon | soon |
| **caveman /sdk** | | | |
| Content-aware compression | Yes | Yes | Yes |
| Byte-exact recovery | Yes | Yes | Yes |
| Tool-schema compression | Yes | Yes | Yes |
| **caveman /cloud** | | | |
| Observability | No | Yes | Yes |
| Harness optimization | No | dev | dev |
| Model fit per task | No | Yes | Yes |
| Evals | No | Yes | Yes |
| Agent-ready | No | Yes | Yes |
| **Deployment and control** | | | |
| No account needed | Yes | No | No |
| Hosted at app.caveman.so | No | Yes | Yes |
| Runs in your VPC | No | No | Yes |
| Spend caps and scoped keys | No | Yes | Yes |
| Data retention you set | No | Yes | Yes |
| Audit log export | No | Yes | Yes |
| License | Apache-2.0 | Commercial | Commercial |

## Published results

| Result | Source |
| --- | --- |
| 33.2% fewer input tokens | [Proxy + Skill study](https://github.com/JuliusBrussee/caveman/blob/main/docs/WRAP-BENCHMARK.md) |
| 63.6% fewer output tokens | [Elasticsearch Labs](https://www.elastic.co/search-labs/blog/elastic-caveman-ai-token-reduction) |

Proxy + Skill benchmark: six Claude Code tasks, three paired runs each. Both groups passed 18/18 exact-answer checks. Input-token reduction, not a promised reduction in your bill.

Elasticsearch Labs measured elastic-caveman across eight live MCP scenarios: 1,284 normal output tokens versus 467 Caveman output tokens. This is a separate study and workload from the Proxy + Skill input benchmark.

## Programs

### Design partners

Cloud is in private development. Design partners run it on their own traffic and shape what ships next. Pricing is set with the founder, on the AI spend Caveman manages.

[Book a call](https://cal.com/caveman/chat)

### Open source

The agent and the SDK are Apache-2.0. Use them at work, read every line, fork them. Model-provider charges still apply, on your own keys.

[Browse the code](https://github.com/JuliusBrussee/caveman)

### Honest numbers

Every figure on this page names its source and its workload, including where Caveman saves nothing. Measured, inferred and verified savings stay separate labels.

[Read how we measure](https://caveman.so/news/does-the-caveman-skill-save-money)

## Questions

### Is Caveman really free?

The agent and the SDK are free and open source, with no Caveman account and no credit card. You still pay your model provider for inference, on your own keys, and cover your own infrastructure.

### How are Pay-as-you-go and Enterprise priced?

There are no seats on any plan. Pay-as-you-go is self-serve: add a card and pay for each product's usage above the free allowances, up to billing limits you set. Enterprise is priced in a contract.

### Will my bill drop by these percentages?

Not necessarily. Caveman measured input tokens with Proxy + Skill; Elasticsearch Labs measured output tokens with elastic-caveman. They are separate studies and do not add together. Results vary by workload, and cache pricing, retries and recovery calls affect the final bill.

### Does my code leave my machine?

The agent and the SDK compress locally and keep the original bytes on your machine; the smaller request goes to your model provider as before. Cloud sees the traffic you route through it. A customer-owned install keeps that traffic in your own cloud account. Hosted Enterprise stores it on Caveman Cloud, redacted and encrypted, and an admin can switch storage off.

### What is the difference between the agent and the SDK?

The agent sits around the coding agent you already use, like Claude Code or Codex. The SDK goes inside agents you build, as middleware for LangChain, the Vercel AI SDK and the OpenAI and Anthropic SDKs. Both are free, and they work on their own or together.

### How are the free tools licensed?

The skill, CLI, SDKs, middleware, extension and Browse are Apache-2.0. Releases before 3.0.0 keep the licence they shipped with. Caveman Code and the standalone cavemem package are MIT. Check each repository for the full terms.

### Can Cloud run in our own cloud?

Yes. Enterprise installs one Helm chart and the same signed images as hosted Cloud on your Kubernetes, on AWS or GCP. Your keys, KMS and storage stay in your account.

[Explore every product](https://caveman.so/products)
