A report card for AI coding

What you built with AI, what it would have cost to hire out, and how well you worked.

Tally reads the coding-agent sessions already on your machine, names each thing you built, prices it against a freelancer, an employee and an agency, and grades six habits against Anthropic's study of 400,000 sessions.

npx tally-score score
See my card Free, with no sign-up. My 864-prompt history took 136 seconds. Works with Claude Code, Codex CLI, OpenCode and OMP.
My cardJune to September · 314 tasks · $5,420 of AI spend

What I built: 74 things, worth $78k–$167k to hire out to freelancers

Contractor outreach and reply pipeline$6.4k–$13k · AI $150
Speculative-decoding fine-tune benchmark$5.6k–$11k · AI $331
Difficulty-tiered BAML pipeline$4.5k–$8.9k · AI $127
Hosted token price and revenue tracker$3.9k–$8.7k · AI $118
CVerificationDoes work finish with proof?22% verified
ADelegationHow much gets done per prompt?15 actions per prompt
BHygiene and recoveryHow much leaks, and is troubled work rescued?15% leaked · 34% recovered
CSpecificationDo requests say what, where, and what done looks like?1.0 of 3 · 5% corrections
CModel disciplineIs routine work on frontier models?22% routable
ALeverageHow much value does a dollar of tokens return?$25 back per $1

Why I built it

I spent $5,420 on AI coding in three months and could not say what it bought.

So I wrote a tool to read my own sessions. It found 74 things I had built. Hired out to freelancers, they would have cost $78,000 to $167,000, and I built them with $4,583 of that spend. Twenty-five had shipped. The rest were experiments, like an NFL simulator meant to beat the Vegas lines. It never did, and it turned into two other projects.

Then it graded how I worked. Only 22% of my tasks ended with proof they worked: a passing test, a commit or my own say-so. Anthropic's study puts novices at 15% and experts at 33%. I had got lazy, and the number said so before I did.

What you built

Each thing you built, priced three ways.

Tally groups your prompts into tasks and your tasks into the things you would name: a pipeline, a report, a fix. Each one is priced at the hours a person in the right role would have needed, as a freelance quote, an employee's loaded cost and an agency rate, with the AI spend on the same scale.

One shipped deliverable from my card, a difficulty-tiered pipeline with a plan-review gate: about 66 hours of senior-engineer work, built with $127 of tokens, or under $2 for each hour replaced. The AI spend is exact, from token usage at published rates. The 66 hours is a model's estimate from the sessions; on my data, five re-runs agreed on 90 to 95% of its judgments. Value means labour replaced, not revenue.

How it works

One command. Stripped before it leaves. Graded in about two minutes.

01

Run one command

npx tally-score score

It reads the session logs your coding agent already keeps. To see what it found without sending anything, add --scan-only.

02

Stripped on your machine

API keys, tokens, emails, phone numbers and card numbers become placeholders before anything is uploaded, and raw transcripts and source files never leave. People's names are removed on receipt, before analysis. Add --metadata-only to send counts and timing with no text at all.

03

Read your card

The card prints in your terminal: what you built, six grades, and the two habits that would move them. To see it in a browser, run npx tally-score claim and add an email.

Claude Code · tested on real sessionsOMP · tested on real sessionsCodex CLI · tested on sample sessionsOpenCode · tested on sample sessions

Six dimensions

Six grades, each answering one question about your sessions.

Each dimension answers one question from the evidence in your sessions, and the cut points are the same for everyone. Three come straight from Anthropic's published numbers. Specification rests on a published finding but uses our thresholds, and model discipline and leverage are our own defaults. Each card says which is which.

C

Verification Anthropic's numbers

Does work finish with proof, not just a judgment?

A ≥ 33% of tasks verified by tests, a commit or your confirmation  ·  B ≥ 25%  ·  C ≥ 15%

Ask for the check in the brief: tests must pass, show the diff, compare against the old output. Novices verify 15% of sessions and experts 33%.

A

Delegation Anthropic's numbers

How much does the agent get done per prompt, and is big work right-sized?

A ≥ 12 actions per prompt  ·  B ≥ 8  ·  C ≥ 5  ·  large tasks that keep being abandoned lower it

Delegate whole features, investigations and documents in one brief, then let the agent run. Novices average about 5 actions per prompt and experts about 12.

B

Hygiene and recovery Anthropic's numbers

How much spend leaks, and does troubled work get rescued?

A ≤ 5% leaked and ≤ 7% abandoned  ·  recovery ≥ 60% of troubled tasks

When a task hits an error, stay with it and say what to check next. Novices abandon 19% of troubled sessions and experts 5 to 7%.

C

Specification Our thresholds

How well do requests say what, where, and what done looks like?

A ≥ 2.2 of 3 on the rubric  ·  corrections ≤ 5%  ·  plans first ≥ 30% of tasks

State the goal, the files in scope, the constraints, and how you'll judge it done. Ask for a plan first on anything large.

C

Model discipline Our default

Is routine work on frontier models, and is oversight spent where it matters?

A ≤ 8% of spend routable to a cheaper model  ·  B ≤ 15%  ·  C ≤ 25%

Set subagents and well-specified chores to a cheaper model; auto-approve low-risk work and interrupt on exceptions.

A

Leverage Our default

How much value does a dollar of tokens return?

A ≥ $20 back per $1  ·  B ≥ $10  ·  C ≥ $4

Put AI on bigger, checkable deliverables; small chat-style asks cost nearly as much per prompt and return far less.

The overall grade is the average of the six on a four-point scale: A at 3.5 and above, B at 2.75, C at 2.

The research behind the grades

Three cut points come from Anthropic's study of 400,000 Claude Code sessions.

Where a published number exists, it sets the line between grades. Where none does, the line is ours, printed on the card, and it will give way to percentiles once 50 people have run Tally on a dimension.

Agents shift worker effort from implementation to supervision, which especially benefits verifiable work and expert workers.
Suproteem K. Sarkar, University of Chicago, AI Agents and Higher-Order Work (2026), from Cursor usage data

Delegation

About five agent actions per prompt in novice sessions and about twelve in expert sessions, across every kind of work.

Anthropic, Agentic coding and returns to expertise

Verification

Success is split into judged and verified, where verified means passing tests, a commit or explicit confirmation. Verified success runs about 15% for novices and 33% for experts, which is where our C and A lines sit.

Anthropic, Agentic coding and returns to expertise

Hygiene and recovery

Novices abandon about 19% of troubled sessions and experts 5 to 7%, and experts rescue troubled sessions to a verified result several times as often.

Anthropic, Agentic coding and returns to expertise

Specification

Experienced Cursor users ask fewer questions and are more likely to set out a plan in the first message. The finding is published; the 0-to-3 rubric and its thresholds are ours.

Sarkar, AI Agents and Higher-Order Work

Model discipline

Experienced users auto-approve more and interrupt more: oversight by exception rather than approval of every step. We show both rates beside the grade; the routing cut points are ours.

Anthropic, Measuring agent autonomy

Value

Anthropic's internal study found 27% of AI-assisted work would not have been done otherwise, so Tally reports labour replaced and work enabled as two numbers, with human-expert time as the unit, as METR does.

Anthropic, How AI is transforming work at Anthropic · METR, time horizons

Also drawn on: Baumann et al., SWE-chat. Across 6,000 public agent sessions, users pushed back in 44% of turns and 44% of agent code survived to a commit, which is why corrections are measured and code survival is next on the list.

Private by default

Tally grades how well you spend, and the grade stays yours.

Leaderboards already rank developers by spend; Viberank ranks about 1,200 of them. Inside companies they have gone badly. Meta's internal token dashboard came down two days after it leaked to the press (Fortune), and Amazon shut down its KiroRank leaderboard after its SVP told staff: "Please don't use AI just for the sake of using AI" (Yahoo Finance).

Your card sits under an anonymous account whose key lives on your machine. When your company uses Tally, it sees team totals, and your own grade stays yours unless you choose to share it.

What leaders are saying

What engineering and finance leaders told TechCrunch in June.

Three worries recur: leaders cannot see the spend, cannot prove the value, and watch the bill outrun the budget.

One of my engineers spent $40,000 on tokens last month, and I genuinely don't know whether I should stop him or should I go and tell everyone else to be like him.
Vitaly Gordon, CEO, Faros AI · TechCrunch, June 2026
In April and May, I started hearing from companies: ‘Oh my god, we are 3x over our entire 2026 token budget and it's only April.’
J.R. Storment, Executive Director, FinOps Foundation · TechCrunch, June 2026
Now the conversations are about, ‘hey, we're spending so much. What visibility do you have? What auditability do you have? What token controls do you have? What is the efficiency of your models?’
Alexander Embiricos, Head of Enterprise, OpenAI · TechCrunch, June 2026
The best ROI comes from moving the broad middle from low to moderate usage, not pushing heavy users higher.
Nicholas Arcolano, Head of Research, Jellyfish · TechCrunch, June 2026

For teams

For a team: what the spend built, what a cheaper model would have saved, and which habits to coach.

43% of 101 tech leaders surveyed by Retool were over their 2026 AI budget. Tally gives a team the ledger of what its spend built, the saving from routing routine work to a cheaper model, and the habits worth coaching, reported as team totals.

Deliverables priced three ways

Every shipped outcome next to its freelance range, internal loaded cost and agency rate, with a frozen baseline and period-over-period change on every number.

Cheaper-model routing savings

Which prompts and delegated runs a cheaper model would have handled, and what that saves. On my account it was $760 of $5,420. A simulator moves the threshold and shows the number change.

Coaching without a leaderboard

The two habits that would move the team's grade, reported for the team. Each person sees their own card and decides whether to share it.

Budget pacing

Spend to date, run-rate, projected month end, days of runway, and the days that broke from a project's baseline.

Chargeback export

A CSV by cost centre, project and month that finance can book, with deliverables tied to tickets where a tracker exists.

Every agent, one scale

Claude Code, Codex CLI, OpenCode and OMP are read into the same shape, so grades and spend compare across tools. Each tool maker grades only its own.

Free

for individuals: your own card, as often as you like.

Pilot

for companies: run on your team's sessions, priced once we know the team's size and spend.

Privacy, briefly

Keys and identifiers are stripped before anything leaves your machine.

What leaves your machine

  • Each prompt and the agent's reply, with secrets and identifiers replaced
  • Token counts, model names and tool names
  • File names and timestamps
  • The machine account name, so the card has an owner
  • One anonymous count per scan, with no text and no account; --no-count turns it off

What never leaves

  • Raw transcripts
  • Source code files
  • Secrets and credentials
  • Emails, phone numbers, card numbers, government IDs and IP addresses, replaced before upload

People's names are removed on receipt, before analysis, and the stored upload is overwritten with the redacted version. Redaction works from patterns and a model's judgment, so an unusual name or key format can slip through; --metadata-only sends no text if that matters to you. npx tally-score delete removes your account and everything stored, at once. The full policy, in plain English: Privacy policy.

Questions

What people ask before they run it.

Is it accurate?

The cost is exact: every prompt is priced from its token usage at the vendor's published rates. The value is an estimate: a model reads each task and judges the hours a person in the right role would have needed, and the card shows that as a range. On my own data, five independent re-runs agreed on 90 to 95% of task judgments. Measured productivity gains from AI coding tools run 5 to 15% (DX), not the marketed multiples, and I would rather be believed than impressive.

What do you keep?

The redacted episodes you upload (prompt text with placeholders, token counts, model and tool names, file names, timestamps, your machine account name) and what Tally computes from them: tasks, deliverables, judgments, grades and reports. Raw transcripts, source code and secrets never reach us, so we cannot keep them.

Can I delete my data?

Yes. npx tally-score delete removes your account, every upload and every report at once. De-identified statistics pooled across many accounts, such as the median correction rate, are kept and cannot be traced back to you.

Do you rank me against other people?

Not publicly. Once 50 people have run Tally, your card will show where you sit on each dimension, privately. A company using Tally sees team totals, not a ranking of its people.

Which tools are supported?

Four. Claude Code and OMP are tested on real sessions; Codex CLI and OpenCode are tested on sample sessions, because I have no real ones yet. Each is read into the same shape, so a grade means the same thing whichever tool you use. If you use something else, tell us on the team page.

What does it cost?

Nothing for individuals, as often as you like. Companies get the team dashboard, pacing, routing savings and chargeback as a paid pilot; request one and we will talk about scope and price.

Where do the grades and cut points come from?

Three cut points use published numbers from Anthropic's study of 400,000 Claude Code sessions: delegation (about 5 actions per prompt for novices, 12 for experts), verification (15% and 33% verified) and abandonment (19% and 5 to 7%). Specification follows a finding from Sarkar's Cursor study with our thresholds, and model discipline and leverage use our own defaults. Every card prints its cut points.

How is value estimated?

For each task: the hours a person would have needed, times the chance they would have done the task at all, times the rate for the role that would have done it, times the chance the work landed. Tasks are then grouped into deliverables and judged once at that level, so hours are not counted twice. The freelance, internal and agency prices are those hours at published rate bands, which a team can replace with its own.