Outcomes

The CFO view: what AI generated this quarter, what it cost, and how confident we are in the numbers. Backed by A/B tests where possible, before/after comparisons elsewhere. Every dollar here ties back to a verified process metric.

Q1 2026 · AS OF MAR 21, 2026
$4.7Mnet value
$4.7M generated, $640K spent — 7.3× return on AI investment
HIGH CONFIDENCEBacked by 4 A/B tests, 12 process metrics, and verified time-tracking samples.
GROSS VALUE
$5.4M
+42% QoQ
TOTAL COST
$640K
+18% QoQ
RETURN ON AI
7.3×
per dollar invested
VALUE GENERATED · $5.4M
Hours saved$2.1M
41,600 hrs × $50 blended rate
Revenue assisted$1.8M
Sales-cycle accel + outbound lift
Costs avoided$980K
Headcount deferral + churn save
Errors prevented$460K
Incident downtime + rework
COST INCURRED · $640K
Model spend$384K
Claude / GPT / inference
Proxon platform$120K
License + storage + audit
Build & maintenance$96K
0.6 FTE allocated
External tools$40K
MCP servers, vendors

Value vs. cost · 8 quarters

VALUECOST

Outcome by process

Where the value is coming from — broken down by the business process AI is touching. Each card shows the before/after on the metrics that matter to the function owner.

CUSTOMER OPS

Customer support

47% of tickets resolved by AI
$1.8M
QUARTERLY VALUE
HIGH
Tickets resolved
18,400/qtr
Avg handle time
14.2min6.8min
CSAT
7.1/107.4/10
Escalation rate
18%12%
AGENTS
tier1-handler-v3ticket-routerknowledge-search
SALES

Sales outreach

$1.2M attributed pipeline
$1.2M
QUARTERLY VALUE
MEDIUM
Emails drafted
4,200/qtr
Reply rate
8%12%
Meetings booked
240/qtr380/qtr
Pipeline velocity
42days31days
AGENTS
drafts_personalized_outbound_emailslead-scorer-42meeting-prep
ENGINEERING

Engineering productivity

12,400 hours saved across 84 engineers
$620K
QUARTERLY VALUE
MEDIUM
PRs reviewed by AI
2,840/qtr
Time-to-first-review
4.2hrs1.1hrs
Standup digests
1,120/qtr
Auto-tagged issues
8,400/qtr
AGENTS
pr-reviewerauto-tags-issues-97digest_eng_standups_v2
FINANCE

Finance & FP&A

Closed the books 4 days faster
$340K
QUARTERLY VALUE
HIGH
Reconciliations
12days4days
Variance explanations
280/qtr
Forecast accuracy
88%94%
AGENTS
reconciles-stripe-vs-erpvariance-explainerforecast-builder
MARKETING

Content marketing

3.2× content output, same headcount
$280K
QUARTERLY VALUE
MEDIUM
Blog posts shipped
24/qtr78/qtr
Avg time per post
16hrs6hrs
Organic traffic lift
0%28%
AGENTS
draft-blog-postseo-optimizersocial-repurpose
LEGAL

Contract review

First-pass review in minutes, not days
$220K
QUARTERLY VALUE
HIGH
Contracts reviewed
84/qtr312/qtr
First-pass time
3.2days0.4days
Issues caught
4.8/contract7.2/contract
AGENTS
contract-redlinerclause-extractor

Adoption vs. impact

Every active agent plotted by adoption (% of target users) and outcome impact ($K/mo). The top-left quadrant is the next opportunity — high impact, low adoption. Bottom-right is engagement traps: people use them but they don't move metrics.

CUST OPSSALESENGFINANCELEGALMKTGPRODUCT

Estimated Time Saved

Rough estimate of human-hours reclaimed in the last 30 days, based on prompt count × average task-time-saved per agent type. · DIRECTIONAL, NOT AUDITED

HOURS RECLAIMED · 30D
7,328hrs
≈ 46 FTEs of capacity
EST. PAYROLL VALUE
$400,000
Blended $50/hr · weighted by dept
VS. LICENSE SPEND
8.4×
Reclaimed value ÷ Proxon spend
Operations
1,842hrs
Sales
1,426hrs
Engineering
1,120hrs
Marketing
924hrs
Finance
768hrs
Product
612hrs
People
384hrs
Legal
252hrs

What's actually attributable

For each major outcome, the actual result vs. the modeled baseline (what we'd expect without AI). A/B tests where we have them, controlled before/after where we don't, flagged estimates where neither was possible.

Actual vs. modeled baseline

What we'd expect to see without the AI tools, based on tests and historical control periods.
Tier-1 tickets resolved
A/B TEST (2 WEEKS, 8 REPS)HIGH
Half of West-coast tier-1 reps had AI assistance, half didn't. Resolution rate gap: 39%.
ACTUAL
18,400 tickets
BASELINE
11,200 tickets
+7,200 tickets (+64%)
Outbound reply rate
A/B TEST (6 WEEKS, 50/50 SPLIT)HIGH
Statistically significant (p=0.003). Effect held across segments.
ACTUAL
12%
BASELINE
8%
+4% (+50%)
Time to first PR review
BEFORE/AFTER, CONTROLLED FOR TEAM SIZEMEDIUM
Q4 2025 baseline measured before pr-reviewer rollout. Team headcount stable.
ACTUAL
1.1 hrs
BASELINE
4.2 hrs
+3.1 hrs (+74%)
Days to close books
BEFORE/AFTER, 6 MONTHS EACHHIGH
Books closed 8 days faster on average since Q3 2025 reconciliation agent went live.
ACTUAL
4 days
BASELINE
12 days
+8 days (+67%)
Content output (blog posts)
ESTIMATE · SAME HEADCOUNT COMPARISONMEDIUM
Marketing team unchanged at 4 FTE. Output up 2.4×. Some lift may be format simplification.
ACTUAL
78 /qtr
BASELINE
32 /qtr
+46 /qtr (+144%)
Pipeline created
ESTIMATE · MULTI-TOUCH ATTRIBUTIONLOW
Multi-touch attribution gives AI tools 18% credit on assisted deals. Confidence interval is wide.
ACTUAL
$8.2M
BASELINE
$6.8M
+$1.4M (+21%)

What got harder

The honest list — second-order effects, drift, and tradeoffs we're tracking.
5 ITEMS
⚠ WARN
+22% escalation handle time
since Jan 2026
Tier-1 escalation timing got worse
Tickets escalated to humans take 22% longer to resolve than before — likely because tier1-handler is auto-resolving the easy cases, leaving harder ones for the queue.
Tune tier1-handler thresholds
⚠ WARN
14 negative tone flags
vs 2 last quarter
Outbound email tone drift
Customers flagged 14 emails as "too generic" or "obviously AI" in Q1 — up from 2 in Q4. drafts_personalized_outbound prompt may need a refresh.
Review prompt + add tone evals
⚠ WARN
3 forecasts off by >10%
this quarter
Forecast over-confidence
forecast-builder shipped 3 forecasts that missed by >10%. Model produces tight confidence intervals even when input data is sparse — finance team learned to discount.
Add uncertainty bands
ℹ INFO
+47% review time on AI PRs
Q1 measurement
Engineers spending more time on AI-generated PRs
Time-on-PR for engineers reviewing AI-drafted code is higher than reviewing human PRs (4.1h vs 2.8h). Net is still positive — drafts come faster — but review fatigue is real.
Improve PR draft quality eval
ℹ INFO
17 shadow assets
+13 vs Q4
Shadow tooling grew alongside official adoption
17 unmanaged tools/models detected — up 13 from last quarter. People are reaching for whatever works, even when registered alternatives exist.
Faster registration workflow