---
title: "Caveman Cloud"
description: "See what your agents cost. Cut it, with proof."
canonical: https://caveman.so/cloud
last-updated: 2026-10-07
---

# Caveman Cloud

See what your agents cost. Cut it, with proof.

Caveman Cloud records the coding agents your team runs and the agents you ship,
and ties what they cost to the work they did. It finds waste, puts a daily price
on it, proposes a change with a way to test it, and counts a saving only once
production traffic proves it. Private preview: book with Julius to talk through
access and pricing for your workload.

## Works with

- **Coding agents your team runs.** Claude Code, Codex, Gemini CLI, OpenCode.
  Spend per person, per agent and per merged change.
- **Agents you ship.** Point an OpenAI SDK, Anthropic SDK, Vercel AI SDK or
  LangChain client at Caveman and keep your provider keys. Every call becomes a
  trace.
- **The gateway you already run.** Keep LiteLLM. Its OpenTelemetry exporter
  sends traces to Caveman, directly or through your OpenTelemetry Collector, and
  provider credentials stay in LiteLLM.

## Observe

Spend per developer, team, agent and workflow. Every request with its status,
cost, latency and tokens. Coding sessions link to their branch and to the change
that merged, so each pull request shows the sessions behind it and what they
spent.

## Prove

Every figure says how it was earned:

- **Measured:** spend at catalog prices, failures and retries included.
- **Inferred:** a finding's estimated saving, with a daily range, a sample size
  and a confidence.
- **Tested:** a proposed change replayed against recorded work and your scorers.
- **Verified:** savings counted from production traffic, day by day, backed by
  provider data.

## Optimize

Caveman reads traffic for repeated work, bloated context and the wrong model for
the job. Each finding becomes a brief: what was found, the proposed change, and
how to verify it. Model routing moves one workload to another model after
checking it on saved traffic. Your team decides what ships.

## Govern

Soft and hard caps per project, with an Inbox item and an email when one is
crossed. Rate limits per key and per workflow. Guardrails that mask or block
sensitive content in prompts and responses. An audit log of changes to budgets,
guardrails, keys and roles.

## What stays yours

- Coding-agent traffic goes straight to your provider with your key. Caveman
  gets a routing ask, not the request.
- On Enterprise, Cloud runs in your VPC, on your KMS and storage, with content
  sharing locked off.
- The local tools need no account and keep working if Cloud is unreachable.

## Links

- Interactive demo (fictional data): https://caveman.so/demo
- Book with Julius: https://cal.com/caveman/chat
- LiteLLM: https://caveman.so/solutions/litellm
- Enterprise: https://caveman.so/enterprise
