Should You Measure Individual Developer Performance?
NOT THE WAY YOU'RE ABOUT TO
Hover or focus to flip ↻The honest answer is: not the way you're about to. What per-person developer metrics destroy, and which questions about people you can answer safely.
TL;DR: You can't rank your way to a better team. Individual activity metrics (commits, lines, PR counts) measure motion, get gamed the day they become targets, and burn the trust you need for everything else. But "never look at people data" is also wrong: knowing who's overloaded, who's a silo of one, and who's carrying invisible load is a leadership duty. The line that matters isn't individual vs. team. It's support vs. ranking.
Every engineering leader eventually gets the email. Board member, CFO, sometimes a founder: "Can we see performance by developer?" It sounds reasonable. You have the data. Git remembers everything. Surely somewhere in those commits is a spreadsheet that says who's great.
Here's the honest answer, and it has two halves that most takes flatten into one: the ranking you're imagining will damage your team, and there are questions about individual people you genuinely owe it to them to answer. Let's do both halves properly. (If the question behind the question is really "is the org productive?", that has its own honest measurement setup, and it isn't per-person either.)
Why individual developer performance metrics are so tempting
Engineering is expensive, opaque, and the biggest line on most software companies' P&L. Leaders outside the codebase can't see the work, so they reach for what's countable: commits, pull requests, story points, lines of code. The data is right there. The dashboard practically builds itself.
The problem is that what's countable and what's valuable barely overlap. The SPACE framework, the closest thing this field has to a measurement standard, is unambiguous: developer productivity "cannot be measured by a single metric or dimension," and activity counts are explicitly its weakest dimension. A staff engineer who spent the week unblocking three people, killing a bad design in review, and deleting 4,000 lines of dead code shows up in an activity dashboard as the team's worst performer. She was its most valuable.
What individual performance metrics actually destroy
The data itself. Developers are professional optimizers - that's the job description. Score commits and commits get smaller and emptier. Score PR throughput and reviews become rubber stamps. Score lines of code and nobody ever deletes anything again. Every activity metric is gameable, and gaming it is the rational response to being scored on it. Within a quarter your dashboard measures compliance theater, not work.
The behavior you needed most. The highest-leverage work on a team (mentoring, glue work, reviewing carefully, writing docs, helping the new hire) produces almost no individual metric exhaust. Rank people on activity and you've made helping a colleague a career mistake. Congratulations: you've priced collaboration out of your own team.
Trust, which was the expensive part. The 2024 Stack Overflow survey is a portrait of what actually blocks developers: 63% name technical debt their top frustration, and 45% agree knowledge silos prevent them from getting ideas across the organization. Those are system failures. When leadership responds to system failures by grading individuals, every developer hears the real message - we will not fix the machine, but we will score you on running in it - and your best people, the ones with options, update their resumes first.


The questions about people that are legitimate
Refusing to rank people is not the same as refusing to see them. A leader who can't answer these questions isn't being ethical. They're being negligent:
- Who is the only person who understands this system? That's key-person risk, and not knowing it is how orgs get held hostage by a resignation.
- Who is overloaded? Work concentrates on the willing. The person carrying a disproportionate share isn't your rockstar. They're your next burnout case and your biggest bus-factor hole, simultaneously.
- Who is doing invisible work the team runs on? Unmeasured glue work is real work. Someone should notice before the promotion committee doesn't.
- Who deserves protection, growth, or more money? Retention is a per-person decision. Pretending otherwise doesn't make you fair; it makes you blind.
Notice the shape of every one of these: the answer is used for the person and the org. Spread the knowledge, rebalance the load, recognize the glue, protect the irreplaceable. None of them requires comparing anyone to anyone.
Scorecards, not stack ranks: the line that holds
The dividing line between legitimate people-visibility and destructive people-scoring:
- Direction of use. A scorecard asks "what does this person uniquely carry, and what do they need?" A ranking asks "who's worse than whom?" The first ends in support; the second ends in a forced curve and a fear culture.
- No single composite number. The moment per-person data collapses into one sortable score, it is a leaderboard, whatever you call it in the deck.
- Receipts, not vibes, and not verdicts either. Evidence-linked observations ("she's the sole reviewer for the payments area") beat both gut-feel reviews and metric verdicts. Data raises the question; a human conversation answers it.
- Multiple dimensions, always in tension. SPACE's practical advice (measure at least three dimensions, include at least one perceptual measure) applies doubly to anything touching individuals. A number without the person's own account of the work is half a fact.


How to answer the board's question
So when the email arrives asking for performance-by-developer, here's the reply that's both honest and useful: "Here's our delivery system's health: flow, queues, knowledge distribution, where work waits. Here's what each person uniquely contributes and what we'd lose without them. And here's why I won't give you a ranked list: because the week I build one, the data starts lying and the best people start leaving - and then we'd deserve the dashboard we'd have."
Measure the system hard. See the people clearly. Never confuse the two. That's the entire discipline of engineering team health, and it fits in one sentence: metrics for systems, judgment for humans, ranking for neither.
Frequently asked
Should you measure individual developer performance?
Not with activity metrics, and never as a ranking. Commit counts, lines of code, and PR tallies measure motion, not value, and they corrode trust the day they become scores. The legitimate version is a scorecard: understanding what each person uniquely carries, where they're overloaded, and what the org would lose without them - protection and growth, never a leaderboard.
Why do individual developer metrics backfire?
Because every activity count is gameable and developers are professional optimizers. Score commits and you get more, smaller, emptier commits. Score review speed and you get rubber stamps. Meanwhile the SPACE research is blunt: productivity can't be captured by a single metric, and activity is its weakest dimension. You end up with worse data and less trust than you started with.
What questions about individual developers are legitimate?
Who is the only person who knows this system? Who is overloaded and heading for burnout? Who is doing invisible glue work that the team depends on? Who deserves protection, growth, or a raise? These are support questions, not ranking questions - they use per-person data to help the person and de-risk the org, not to grade anyone against a peer.
What should engineering leaders measure instead?
The system: pickup times, queue lengths, knowledge concentration, review load distribution, delivery flow. Measure across multiple SPACE dimensions, include at least one perceptual measure, and aggregate at team level. When the system metrics are healthy and the team says the work feels good, individual performance mostly takes care of itself.