Engineering intelligence

Your agents spend.
What ships?

Connect coding-agent sessions to the work your team delivers. See model spend per PR, across repositories, and by the kind of work that lands.

Caveman Platform · private development

Coding sessionsRepository contextDelivered work
See the whole picture

The bill counts tokens. Your team ships software.

A million tokens.
How much progress?

Every PR.
Its cost to land.

Explore the work behind a sample team's model spend. Filter features, fixes, and docs. Select a pull request to see its sessions, author, and attributed cost.

acme / productDelivery intelligence · example data
Cost to merge · attributed PRs$57.04
Merged PRs08
Attribution coverage88%
Cost per pull requestClick a point to inspect
$0$5$10$15$200h8h16h24h32h#241#242#243#244#245#246#248
Time from open to merge →
Interactive sample · catalog-priced model usage, not labor cost or a productivity score · unlinked PRs excluded from cost totals

PR attribution uses matching branch tags. Work types follow branch prefixes or conventional-commit titles. Missing links remain unknown; cost alone is not a measure of developer performance.

Your existing agents.
A shared picture.

Keep the coding tools your team already uses. Bring supported sessions, usage records, and repository context into the same workspace.

01

Connect your agents

Collect supported coding sessions and model usage.

02

Bind your repository

Join branch-tagged sessions to merged pull requests.

03

Follow the work

Compare PRs, repos, work types, and attribution coverage.

A team-wide view.
Down to the work.

Move between teammates and work types using the same underlying records. Keep attribution coverage beside the cost so a missing connection cannot distort the comparison.

Team economics / shared workspaceSame sample PR records · grouped views
Attributed model spendAlphabetical · select a group

Shared infrastructure questions. No developer leaderboard.

Sample model costs, not labor costs · group comparisons are descriptive, not causal measures of productivity

Find the pattern.
Not a scapegoat.

A single expensive run tells you little. Compare paths through related work, inspect repeated calls, and find the workflow worth investigating next.

Workload analysis / task familyExample trace patterns
Repeated verification
01Read
02Verify
03Read
04Verify
05Answer

The same artifact gets read and verified twice. Inspect the traces before deciding whether either check is redundant.

Illustrative pattern analysis · themes require suitable captured evidence · lexical clusters are not labeled semantic

Team and repository context.

Follow who ran a session and which repository it belongs to. Keep organization and personal views distinct.

Work that has not landed.

Separate merged, in-flight, and older unlanded branches. Today's active work is not automatically waste.

Interpretation with evidence.

Inspect the traces behind a theme. Semantic analysis depends on suitable content and analysis inputs; PR cost remains based on measured links.

Run receipt
Model calls2
Tool calls3
Stop reasonCompleted
Illustrative receipt

A feature has a paper trail.

Compare feature PRs and repositories using attributed model spend. Multi-PR feature accounting needs explicit grouping, not a guessed semantic match.

Unlisted modelUnpriced

Missing price is not a zero-dollar call.

Know what you are missing.

Keep attribution coverage next to spend. A merged PR without a linked session is unmeasured, never free.

Workflow: repair-tests
JD
Engineering / JulesAgent: code-review

A team view with context.

Understand agent adoption and workflow costs. Use the data to improve the system, not rank people by token usage.

The delivery view.
Built for real questions.

Move from team activity to delivered changes, then inspect the record. These are full screens from a seeded Caveman workspace.

● ● ●
Demo workspace
Caveman delivery analytics in a seeded workspace, with merged pull requests and attributed cost.

Real product UI · seeded demo workspace · catalog list-price subtotals

Does this measure the full cost of a feature?

It measures attributed model usage, not labor or total engineering cost. PR-level links come from branch-tagged sessions. Work-type summaries group feature, fix, and other PR categories; they do not automatically infer a business feature across unrelated PRs.

What if a session has no branch tag?

Its measured spend remains visible without a delivery outcome. Coverage shows how much session traffic can be linked, so incomplete instrumentation cannot look like a free PR.

Do you need to capture every prompt?

Session and branch attribution use metadata. Deeper content analysis has separate capture, consent, and retention requirements. What you connect determines the available evidence.

Make every token count.

Know what AI
helps you ship.

Bring your coding agents, repositories, and team questions. Explore Engineering intelligence in Caveman Platform.