Caveman/data-use← back
Caveman Cloud · what we collect

What we collect

Draft · version 2026-07

Draft — not yet in effect

This document is under legal review and is published here for transparency only. It does not yet bind you or us. We'll set an effective date once review is complete.

A plain-language map of what Caveman collects on each plan, and what we never collect. It's the ten-second version of the Cloud Privacy Policy.

By plan

WhatFreeIndieTeamEnterpriseLocal wrap
Usage metadataspans, tokens, cost, latency, statusYesYesYesYesToken counts only
Payload hashessalted, for dedupe and cache lookupsYesYesOptionalOptionalLocal only
Raw request & response payloadsthe prompt and response bytesYesYesDefault onNeverNever leaves your machine
Used to improve Caveman's modelstraining and evaluationYesYesDefault onNeverNever — no payloads exist to use
Product analyticsroutes, section reads, engagement & feature usage; cookielessYesYesYesNevern/a
CLI telemetryanonymous, no content or pathsOpt-inOpt-inOpt-inOpt-inOpt-in (off by default)
Cross-customer practice evidencederived before/after metric, window sizes, practice idOpt-inOpt-inOpt-inNevern/a

“Default on” means it starts on and you can turn it off any time in Data governance. On Free and Indie, product data sharing is a condition of the plan — it's part of what keeps them cheap. On Enterprise it is off, contractually.

The local wrap is stricter than every hosted plan: its telemetry is token counts, model names, and savings numbers — never prompt or response bytes, which have no upload path at all. On Free and Indie that telemetry is part of the plan (it is also how the weekly allowance is metered); on Team it has an org-level opt-out; on Enterprise zero-data-retention orgs the endpoint refuses the upload entirely.

Your own retention window is what applies. Every organization starts at 30 days for stored payloads, and that is the window we delete on. Payloads covered by product data sharing have a longer ceiling — up to 365 days — but that ceiling only comes into play if you raise your own window above 30 days. The shorter of the two always wins. You can see and change your window in Data governance.

Two things are worth being precise about. First, the last two rows above are two separate switches, and we keep payloads only while both are on. Turning off product data sharing stops new payloads being kept and purges what was already collected under it; turning off payload storage stops it as well. Second, compression recovery storage (the encrypted originals that let caveman restore compressed prompts byte-for-byte) is separate again — it exists to serve you, follows your retention window, and is never used for training. Zero-data-retention mode disables all of it.

Cross-customer practice evidence is separate from product data sharing. It is off by default and requires an explicit project-level opt-in. Only derived before/after metrics, window sizes, and a practice id may contribute — never prompts, responses, diffs, repository content, or organization identity. Revocation applies to the next rollup. Enterprise and organizations without a known plan cannot opt in.

What we never collect

  • The CLI never sends your prompts, completions, code, or file paths — only anonymous command metadata, and only if you opt in.
  • The local wrap never sends prompt or response bytes — there is no code path that uploads them. Its telemetry is token counts and savings numbers only.
  • The browser extension collects nothing. See its own privacy policy.
  • Enterprise traffic is never used for training or product improvement, and product analytics is off for enterprise organizations.
  • We never sell your data, and we never share your prompts or responses with third parties for their own purposes.
  • Product analytics never includes prompt or response content, or any text you typed.