A report card for AI coding
Tally reads the coding-agent sessions already on your machine, names each thing you built, prices it against a freelancer, an employee and an agency, and grades six habits against Anthropic's study of 400,000 sessions.
What I built: 74 things, worth $78k–$167k to hire out to freelancers
Why I built it
So I wrote a tool to read my own sessions. It found 74 things I had built. Hired out to freelancers, they would have cost $78,000 to $167,000, and I built them with $4,583 of that spend. Twenty-five had shipped. The rest were experiments, like an NFL simulator meant to beat the Vegas lines. It never did, and it turned into two other projects.
Then it graded how I worked. Only 22% of my tasks ended with proof they worked: a passing test, a commit or my own say-so. Anthropic's study puts novices at 15% and experts at 33%. I had got lazy, and the number said so before I did.
What you built
Tally groups your prompts into tasks and your tasks into the things you would name: a pipeline, a report, a fix. Each one is priced at the hours a person in the right role would have needed, as a freelance quote, an employee's loaded cost and an agency rate, with the AI spend on the same scale.
| Way to get the same 66 hours of work | Cost |
|---|---|
| AI spend, exact from token usage | $127 |
| Freelance quote range, marketplace rates with the client fee | $4,500 – $8,900 |
| Internal loaded cost for a senior engineer | $10,400 |
| Agency rate | $13,800 |
One shipped deliverable from my card, a difficulty-tiered pipeline with a plan-review gate: about 66 hours of senior-engineer work, built with $127 of tokens, or under $2 for each hour replaced. The AI spend is exact, from token usage at published rates. The 66 hours is a model's estimate from the sessions; on my data, five re-runs agreed on 90 to 95% of its judgments. Value means labour replaced, not revenue.
How it works
It reads the session logs your coding agent already keeps. To see what it found without sending anything, add --scan-only.
API keys, tokens, emails, phone numbers and card numbers become placeholders before anything is uploaded, and raw transcripts and source files never leave. People's names are removed on receipt, before analysis. Add --metadata-only to send counts and timing with no text at all.
The card prints in your terminal: what you built, six grades, and the two habits that would move them. To see it in a browser, run npx tally-score claim and add an email.
Six dimensions
Each dimension answers one question from the evidence in your sessions, and the cut points are the same for everyone. Three come straight from Anthropic's published numbers. Specification rests on a published finding but uses our thresholds, and model discipline and leverage are our own defaults. Each card says which is which.
Does work finish with proof, not just a judgment?
A ≥ 33% of tasks verified by tests, a commit or your confirmation · B ≥ 25% · C ≥ 15%
Ask for the check in the brief: tests must pass, show the diff, compare against the old output. Novices verify 15% of sessions and experts 33%.
How much does the agent get done per prompt, and is big work right-sized?
A ≥ 12 actions per prompt · B ≥ 8 · C ≥ 5 · large tasks that keep being abandoned lower it
Delegate whole features, investigations and documents in one brief, then let the agent run. Novices average about 5 actions per prompt and experts about 12.
How much spend leaks, and does troubled work get rescued?
A ≤ 5% leaked and ≤ 7% abandoned · recovery ≥ 60% of troubled tasks
When a task hits an error, stay with it and say what to check next. Novices abandon 19% of troubled sessions and experts 5 to 7%.
How well do requests say what, where, and what done looks like?
A ≥ 2.2 of 3 on the rubric · corrections ≤ 5% · plans first ≥ 30% of tasks
State the goal, the files in scope, the constraints, and how you'll judge it done. Ask for a plan first on anything large.
Is routine work on frontier models, and is oversight spent where it matters?
A ≤ 8% of spend routable to a cheaper model · B ≤ 15% · C ≤ 25%
Set subagents and well-specified chores to a cheaper model; auto-approve low-risk work and interrupt on exceptions.
How much value does a dollar of tokens return?
A ≥ $20 back per $1 · B ≥ $10 · C ≥ $4
Put AI on bigger, checkable deliverables; small chat-style asks cost nearly as much per prompt and return far less.
The overall grade is the average of the six on a four-point scale: A at 3.5 and above, B at 2.75, C at 2.
The research behind the grades
Where a published number exists, it sets the line between grades. Where none does, the line is ours, printed on the card, and it will give way to percentiles once 50 people have run Tally on a dimension.
Agents shift worker effort from implementation to supervision, which especially benefits verifiable work and expert workers.Suproteem K. Sarkar, University of Chicago, AI Agents and Higher-Order Work (2026), from Cursor usage data
About five agent actions per prompt in novice sessions and about twelve in expert sessions, across every kind of work.
Anthropic, Agentic coding and returns to expertiseSuccess is split into judged and verified, where verified means passing tests, a commit or explicit confirmation. Verified success runs about 15% for novices and 33% for experts, which is where our C and A lines sit.
Anthropic, Agentic coding and returns to expertiseNovices abandon about 19% of troubled sessions and experts 5 to 7%, and experts rescue troubled sessions to a verified result several times as often.
Anthropic, Agentic coding and returns to expertiseExperienced Cursor users ask fewer questions and are more likely to set out a plan in the first message. The finding is published; the 0-to-3 rubric and its thresholds are ours.
Sarkar, AI Agents and Higher-Order WorkExperienced users auto-approve more and interrupt more: oversight by exception rather than approval of every step. We show both rates beside the grade; the routing cut points are ours.
Anthropic, Measuring agent autonomyAnthropic's internal study found 27% of AI-assisted work would not have been done otherwise, so Tally reports labour replaced and work enabled as two numbers, with human-expert time as the unit, as METR does.
Anthropic, How AI is transforming work at Anthropic · METR, time horizonsAlso drawn on: Baumann et al., SWE-chat. Across 6,000 public agent sessions, users pushed back in 44% of turns and 44% of agent code survived to a commit, which is why corrections are measured and code survival is next on the list.
Private by default
Leaderboards already rank developers by spend; Viberank ranks about 1,200 of them. Inside companies they have gone badly. Meta's internal token dashboard came down two days after it leaked to the press (Fortune), and Amazon shut down its KiroRank leaderboard after its SVP told staff: "Please don't use AI just for the sake of using AI" (Yahoo Finance).
Your card sits under an anonymous account whose key lives on your machine. When your company uses Tally, it sees team totals, and your own grade stays yours unless you choose to share it.
What leaders are saying
Three worries recur: leaders cannot see the spend, cannot prove the value, and watch the bill outrun the budget.
One of my engineers spent $40,000 on tokens last month, and I genuinely don't know whether I should stop him or should I go and tell everyone else to be like him.Vitaly Gordon, CEO, Faros AI · TechCrunch, June 2026
In April and May, I started hearing from companies: ‘Oh my god, we are 3x over our entire 2026 token budget and it's only April.’J.R. Storment, Executive Director, FinOps Foundation · TechCrunch, June 2026
Now the conversations are about, ‘hey, we're spending so much. What visibility do you have? What auditability do you have? What token controls do you have? What is the efficiency of your models?’Alexander Embiricos, Head of Enterprise, OpenAI · TechCrunch, June 2026
The best ROI comes from moving the broad middle from low to moderate usage, not pushing heavy users higher.Nicholas Arcolano, Head of Research, Jellyfish · TechCrunch, June 2026
For teams
43% of 101 tech leaders surveyed by Retool were over their 2026 AI budget. Tally gives a team the ledger of what its spend built, the saving from routing routine work to a cheaper model, and the habits worth coaching, reported as team totals.
Every shipped outcome next to its freelance range, internal loaded cost and agency rate, with a frozen baseline and period-over-period change on every number.
Which prompts and delegated runs a cheaper model would have handled, and what that saves. On my account it was $760 of $5,420. A simulator moves the threshold and shows the number change.
The two habits that would move the team's grade, reported for the team. Each person sees their own card and decides whether to share it.
Spend to date, run-rate, projected month end, days of runway, and the days that broke from a project's baseline.
A CSV by cost centre, project and month that finance can book, with deliverables tied to tickets where a tracker exists.
Claude Code, Codex CLI, OpenCode and OMP are read into the same shape, so grades and spend compare across tools. Each tool maker grades only its own.
for individuals: your own card, as often as you like.
for companies: run on your team's sessions, priced once we know the team's size and spend.
Privacy, briefly
--no-count turns it offPeople's names are removed on receipt, before analysis, and the stored upload is overwritten with the redacted version. Redaction works from patterns and a model's judgment, so an unusual name or key format can slip through; --metadata-only sends no text if that matters to you. npx tally-score delete removes your account and everything stored, at once. The full policy, in plain English: Privacy policy.
Questions
The cost is exact: every prompt is priced from its token usage at the vendor's published rates. The value is an estimate: a model reads each task and judges the hours a person in the right role would have needed, and the card shows that as a range. On my own data, five independent re-runs agreed on 90 to 95% of task judgments. Measured productivity gains from AI coding tools run 5 to 15% (DX), not the marketed multiples, and I would rather be believed than impressive.
The redacted episodes you upload (prompt text with placeholders, token counts, model and tool names, file names, timestamps, your machine account name) and what Tally computes from them: tasks, deliverables, judgments, grades and reports. Raw transcripts, source code and secrets never reach us, so we cannot keep them.
Yes. npx tally-score delete removes your account, every upload and every report at once. De-identified statistics pooled across many accounts, such as the median correction rate, are kept and cannot be traced back to you.
Not publicly. Once 50 people have run Tally, your card will show where you sit on each dimension, privately. A company using Tally sees team totals, not a ranking of its people.
Four. Claude Code and OMP are tested on real sessions; Codex CLI and OpenCode are tested on sample sessions, because I have no real ones yet. Each is read into the same shape, so a grade means the same thing whichever tool you use. If you use something else, tell us on the team page.
Nothing for individuals, as often as you like. Companies get the team dashboard, pacing, routing savings and chargeback as a paid pilot; request one and we will talk about scope and price.
Three cut points use published numbers from Anthropic's study of 400,000 Claude Code sessions: delegation (about 5 actions per prompt for novices, 12 for experts), verification (15% and 33% verified) and abandonment (19% and 5 to 7%). Specification follows a finding from Sarkar's Cursor study with our thresholds, and model discipline and leverage use our own defaults. Every card prints its cut points.
For each task: the hours a person would have needed, times the chance they would have done the task at all, times the rate for the role that would have done it, times the chance the work landed. Tasks are then grouped into deliverables and judged once at that level, so hours are not counted twice. The freelance, internal and agency prices are those hours at published rate bands, which a team can replace with its own.