Weekly output
Benchmark lines show peer median and top-decile output.
Measure whether engineering is getting faster in ways that survive review, tests, deploys, and customer use. Output, flow, quality, reliability, and AI leverage stay on the same page so velocity cannot hide risk.
3 things worth a closer look.
Normalized units of expert engineering work per engineer, broken out by feature, bug, and KTLO output.
Benchmark lines show peer median and top-decile output.
Follow work from ready state through code, review, rework, and deploy. This separates real acceleration from work sitting in queues.
Track whether faster output creates defects, review debt, unreliable tests, or deploy risk.
Bug-introducing commits per normalized output unit, with introduced-bug count overlaid.
comments with code or test evidence
requested changes resolved in one pass
merged PRs with linked test proof
PRs under 400 changed lines
| Team | Quality | Bugs / output | Escaped | Flaky tests | Review |
|---|---|---|---|---|---|
| Platform Runtime | 86 | 0.018 | 2 | 4 | 62 |
| Product Growth | 89 | 0.016 | 1 | 3 | 71 |
| Customer Automations | 63 | 0.037 | 6 | 12 | 57 |
| Data Apps | 78 | 0.022 | 3 | 7 | 68 |
| Mobile Core | 82 | 0.019 | 2 | 5 | 81 |
| Security Engineering | 91 | 0.012 | 0 | 2 | 76 |
Separate human-only work from assisted and agent-drafted work, then score the downstream review and cost profile.
| Tool | Sessions | Output | Review | Cost / unit |
|---|---|---|---|---|
| Codex | 426 | 31.4 | 66 | $0.42 |
| Copilot | 812 | 28.1 | 72 | $0.18 |
| Cursor | 266 | 21.7 | 61 | $0.31 |
| Claude Code | 188 | 14.9 | 58 | $0.57 |
Velocity gets useful when throughput, flow, quality, reliability, people load, and AI leverage are all visible together.