---
title: "What we collect"
description: "A plain-language summary of what Caveman collects on each Caveman Cloud plan and from the CLI and local tools, how long we keep it, and what we never collect."
canonical: https://caveman.so/data-use
last-updated: 2026-10-08
status: "in effect"
version: "2026-10-08.2"
effective: "October 8, 2026"
publisher: "Caveman Labs, Inc."
contact: "contact@caveman.so"
---

# What we collect

A plain-language summary of what Caveman collects on each Caveman Cloud plan and from the CLI and local tools, how long we keep it, and what we never collect.

This page summarises the [Privacy Policy](/privacy). The Privacy Policy, the [Terms of Service](/terms) and the [Data Processing Agreement](/legal/dpa) govern.

## Where your deployment stores data

Deployment location and subscription plan are separate. Keeping data for replay and evaluation does not mean sending it to Caveman's hosted service.

| Deployment | Request content and operational telemetry | What reaches Caveman outside the deployment |
|---|---|---|
| **Hosted Cloud** | Stored in the infrastructure Caveman operates for the hosted service, under the settings below | Customer Data needed to provide the hosted service |
| **Enterprise in your cloud account** | Request bodies, traces, artifacts and operational telemetry are stored in your deployment's configured databases, object storage and backups | Separate CLI usage telemetry unless disabled; information you provide for support, billing or the customer relationship; any other access or transfer expressly covered by your deployment agreement |

Enterprise in your cloud account retains data to operate the platform. It is not a zero-data-retention service. Your agreement identifies the cloud account and region, who operates the deployment, Caveman's permitted access, and the retention and deletion responsibilities. Your chosen model providers still receive requests and replays according to your configuration. Connecting a client to Caveman's hosted gateway or hosted routing is a separate use of the hosted service.

The CLI's product telemetry is separate from request traces stored in your deployment. It sends usage events and a random install ID to Caveman; the receiving service also records your IP address. It does not send prompts, responses, code or file paths. Disable it with `caveman telemetry off`, `CAVEMAN_TELEMETRY=0` or `DO_NOT_TRACK=1`.

## Hosted plans and local tools

The plan columns below describe hosted Workspaces, including Enterprise where a contract provides for hosting by Caveman. They do not relocate data from a customer-account deployment into Caveman's hosted service.

| What | Free | Pay-as-you-go | Hosted Enterprise | CLI and local tools |
|---|---|---|---|---|
| **Usage Metadata**: model, provider, token counts, cost, latency, status, tags, salted hashes of Payloads | Recorded for every request within your allowances | Recorded for every request within your billing limits | Recorded for every request | Span metadata only, when you are signed in and sync, or in managed gateway mode |
| **Request and response bodies (Payloads)** | Stored by default, redacted and encrypted. Storage cannot be turned off | Stored by default, redacted and encrypted. An Admin can turn storage off | Stored by default, redacted and encrypted. An Admin can turn storage off | Never uploaded, except requests you route through the Caveman Cloud gateway and, with routing on, the routing requests the Privacy Policy describes |
| **Compression originals**: exact copies that let a compressed prompt be restored | Kept encrypted, not redacted. Follow the payload setting | Kept encrypted, not redacted. Follow the payload setting | Kept encrypted, not redacted. Follow the payload setting | Kept on your machine, in `~/.caveman/ccr.db` |
| **Replay to your Model Providers**: redacted stored Payloads, sent on your stored keys | On by default for each project. An Admin can turn it off | On by default for each project. An Admin can turn it off | On by default for each project. An Admin can turn it off | Not applicable |
| **Used to improve Caveman's models** (Product Data Sharing) | Yes, a condition of the plan, agreed on its own screen at sign-up | Off by default for new Workspaces; Workspaces moved from the former Indie or Team plans kept their setting. An Admin can turn it on or off | Never. It cannot be turned on | Telemetry, never. Gateway traffic follows your Workspace's plan |
| **Zero data retention per request**: `x-cave-retention: zdr` | Available on every request | Available on every request | Available on every request; an Order Form can also turn payload storage and artifact storage off for the Workspace | Not applicable |
| **Product analytics in the dashboard** | Not switched on | Not switched on | Not switched on. It will stay off for Enterprise | Not applicable |
| **CLI telemetry**: pseudonymous, with a random install ID and IP address | On by default, opt-out | On by default, opt-out | On by default, opt-out | On by default, opt-out |
| **Cross-customer practice evidence**: derived before-and-after metrics only | Opt-in per project, off by default | Opt-in per project, off by default | Not available | Not applicable |

- **Payloads.** Before a Payload is stored, automated redaction removes common sensitive patterns, such as email addresses, bearer tokens and payment card numbers. Redaction will not catch everything. Payloads are encrypted with AES-256-GCM under data keys scoped to their Workspace.
- **Replay.** With replay on, Caveman Cloud can send a project's redacted stored Payloads to your Model Providers, on your keys, to compare models, build and test fixes, and judge results. Your Model Providers charge you for this. `x-cave-retention: zdr` excludes a request.
- **Product Data Sharing.** It covers what reaches Caveman Cloud from the Workspace, including recorded requests with their Payloads. On Free, for projects that opt in to router learning, it also covers sampled routing-request text. Turning it off deletes the stored copies kept only for it. Training sets already built from it are not rebuilt with it and expire at most 365 days after they were built. We may use data for this purpose only where Product Data Sharing applies and the required notices and permissions are in place.
- **Zero data retention.** `zdr` turns off payload storage, compression originals, caching and replay for that request. Usage Metadata is still recorded.
- **Product analytics.** caveman.so runs its own cookieless analytics, separate from the dashboard. The [Privacy Policy](/privacy) describes it.
- **CLI telemetry.** It is the same on every plan. In CI the CLI sends nothing unless you set `CAVEMAN_TELEMETRY=1`. Turn it off with `caveman telemetry off`, `DO_NOT_TRACK=1` or `CAVEMAN_TELEMETRY=0`.
- **Practice evidence.** Only derived before-and-after metrics, window sizes and a practice identifier contribute. Never prompts, responses, diffs, repository content or organisation identity.

## How long we keep it

The hosted retention and backup timings below apply to hosted Cloud. For a customer-account deployment, the deployment agreement and configured storage and backup policies govern. A retention setting does not authorize use of Enterprise data to improve Caveman's models.

- **Customer Data.** Request history, Usage Metadata, Payloads, compression originals, imported content and agent artifacts are kept until you delete them, on every plan. An Admin can set a window in Data governance, from 1 to 36,500 days, after which they are deleted automatically. By default no window is set. Agent artifacts are kept only while artifact storage is on. It is on for Workspaces that start on Free, and it stays on when they add a card or move to Enterprise. It is off for Workspaces created directly on Pay-as-you-go or Enterprise, including customer-owned installs. An Admin can turn it on or off on every plan.
- **Deletion.** A deletion you request in Data governance is carried out 30 days after the request. Backups roll off about 35 days after deletion. Rows deleted from our analytics database by a retention window can stay on disk, and in backups taken in that time, for up to about two months in total. If an Admin lengthens or removes the window after rows were deleted, those rows can stay on disk until the database merges them or the Workspace is deleted.
- **Derived data.** Prompt snippets, embeddings and cluster maps: 90 days at most. Semantic-cache samples: 30 days at most. Semantic-cache judgments: 90 days at most. None is kept longer than your window. Cached responses: 5 minutes by default, 24 hours at most.
- **Routing and billing.** Routing decision counts: 13 months. Routing outcomes from Free projects that opt in to router learning: up to 365 days. Daily counts of billable events: for the life of the Workspace.
- **Product Data Sharing.** Data is kept while sharing applies to it, and never longer than your window. Training sets built from it expire at most 365 days after they are built.
- **CLI telemetry.** IP addresses are cleared after 90 days, and events are deleted after 13 months.

## What we never collect

- CLI telemetry never sends prompts, completions, raw command-line arguments, file paths, tool arguments or results, credentials, or rows or files from your local databases.
- Our local tools never upload prompt or response bytes to Caveman, with two exceptions you choose: requests you route through the Caveman Cloud gateway, and routing (in private preview). With routing on, the CLI sends your latest message, the one before it and the end of the agent's last reply to Caveman Cloud to pick a model and effort. The text is read in memory and dropped, unless the Free router-learning conditions in the [Privacy Policy](/privacy) apply.
- The Caveman Mode browser extension collects nothing and makes no network requests.
- Enterprise traffic is never used to improve Caveman's models, and neither is traffic from dedicated or self-hosted deployments.
- We never sell personal data. We never share data from Product Data Sharing with other customers, and never use it to train a third party's general-purpose model.
