GranularityweekDatelast 3 monthsOrgEngineering
ENGINEERING OPERATIONS

Dev Team Velocity

Measure whether engineering is getting faster in ways that survive review, tests, deploys, and customer use. Output, flow, quality, reliability, and AI leverage stay on the same page so velocity cannot hide risk.

Insights

3 things worth a closer look.

Active 3AcknowledgedIgnored
highAI-assisted output per engineer surged 3.2xPlatform Runtime and Product Growth drove the spike; review load is now the constraint.ended 5 days ago
mediumReview queue time is masking cycle-time gainsMedian coding time fell 28%, but first-review wait rose to 18 hours in the same period.active
goodBug-introducing commits fell while output climbedBugs per output improved from 0.037 to 0.021 after test-generation guardrails shipped.updated today
Output / engineer
90.4/wk
+577%97th percentile vs. peer cohort
Lead time
6.3days
-34%idea ready to production deploy
Review turnaround
18hrs
+41%first human review is the bottleneck
Change fail rate
6.8%
-2.1 ptsdeploys causing rollback or incident
Bugs / output
0.021
-43%bug-introducing commits per output unit
AI leverage
42%
+19 ptsmerged diff touched by approved AI tools

Output per engineer

Normalized units of expert engineering work per engineer, broken out by feature, bug, and KTLO output.

Weekly output

Benchmark lines show peer median and top-decile output.

90.4/week97th percentile
FeatureBugKTLOProjectedAI share

Output by team

PER WEEK
PRPlatform Runtime108.4/wk
PGProduct Growth98.2/wk
CACustomer Automations86.4/wk
DAData Apps74.9/wk
MCMobile Core61.8/wk
SESecurity Engineering54.2/wk

Output by type

SHARE
Feature62%Bug23%KTLO15%

Output by project

ALL TEAMS
Proxon MVP34%Proxon Dogfooding39%SDK Documentation16%Billing Modernization7%Platform Resilience4%

Flow efficiency

Follow work from ready state through code, review, rework, and deploy. This separates real acceleration from work sitting in queues.

Coding time
1.9d
-28%
Review wait
18h
+41%
Merge to deploy
11h
-22%
Flow efficiency
42%
+9 pts

Cycle-time funnel

MEDIAN DURATION
2.4d
Ready workissue ready to first commit
1.9d
Active codingfirst commit to PR open
18h
Review queuePR open to first review
9h
Reworkrequested changes to approval
11h
Deploymerge to production

Where time goes

FLOW SPLIT
Active coding29%Review queue18%CI / flaky tests7%Waiting on owner26%Release window20%
14PRs older than 3 daysReviewers: rebalance reviewers
23Branches without PRTeam leads: open or close stale work
18Blocked ticketsProduct: resolve acceptance gaps
41Flaky-test rerunsRelease Eng: quarantine noisy suites

Code quality

Track whether faster output creates defects, review debt, unreliable tests, or deploy risk.

Bugs per output

Bug-introducing commits per normalized output unit, with introduced-bug count overlaid.

0.021-43% vs. March

Review quality

TARGET MARKERS
Review depth74

comments with code or test evidence

Resolution quality57

requested changes resolved in one pass

Test evidence81

merged PRs with linked test proof

PR size hygiene69

PRs under 400 changed lines

Team quality comparison

SORT: QUALITY SCORE
TeamQualityBugs / outputEscapedFlaky testsReview
Platform Runtime860.0182462
Product Growth890.0161371
Customer Automations630.03761257
Data Apps780.0223768
Mobile Core820.0192581
Security Engineering910.0120276

AI leverage without quality drift

Separate human-only work from assisted and agent-drafted work, then score the downstream review and cost profile.

AI-touched merged diff
42%
+19 pts
Agentic PRs merged
118
+64%
AI rework rate
12%
-5 pts
Cost / output unit
$1.84
-18%

Merge mix

BY DIFF ATTRIBUTION
Human only38%AI assisted44%Agent drafted18%

Tools by output quality

LAST 30 DAYS
ToolSessionsOutputReviewCost / unit
Codex42631.466$0.42
Copilot81228.172$0.18
Cursor26621.761$0.31
Claude Code18814.958$0.57

Measurement library

Velocity gets useful when throughput, flow, quality, reliability, people load, and AI leverage are all visible together.

Throughput

normalized output per engineerPRs merged per engineerfeature / bug / KTLO mixplanned work delivereddiff accepted per review hourAI-assisted output lift

Flow

lead time for changescycle timereview waitqueue timeflow efficiencyWIP agebranch ageblocked work

Quality

bugs per outputescaped defectsbug-introducing commitsrework ratetest evidenceflaky testscode churncomplexity drift

Reliability

deployment frequencychange failure raterollback rateMTTRincident countCI pass ratebuild timedeploy batch size

Team Health

review load balancehero riskafter-hours workon-call interruptionfocus-time preservationmanager escalation ageonboarding ramp

AI Leverage

AI-touched diff shareagent success rateaccepted suggestionsprompt cost per outputAI rework rategenerated test sharepolicy-compliant tool use