Measure a Forward Deployed Engineer on three levels:
- the scored roadmap of future AI opportunities the engagement uncovered across the business
- the value of the first system shipped, meaning the workflow metric it moved
- the acceleration of every initiative that follows
Most AI ROI models fail because they only measure the system that shipped and miss the compounding value on either side of it.
Sooner or later the question arrives, usually from finance: what did we get for that? It is a fair question, and AI has a well-earned reputation for struggling to answer it.
This guide gives you a measurement framework that survives a CFO's scrutiny, grounded in what the research actually shows about AI returns, and honest about the costs on the other side of the ledger. If you are trying to justify a Forward Deployed Engineer, or prove the value of one you already engaged, start here.
The short version
Key takeaways
- AI ROI is genuinely hard to measure, and the research shows most organizations struggle: in one large survey, fewer than 1% reported significant returns.
- The usual mistake is measuring only the first shipped system and missing where the real value compounds.
- A three-level model fixes this: value found, value shipped, and value multiplied.
- Set the baseline before the engagement starts, or you will be arguing about attribution afterward.
- The honest cost comparison is not rate versus rate. It is a Forward Deployed Engineer versus the in-house market, and versus another stalled pilot.
Why AI ROI is hard to measure, and why that excuse is expiring
Two things are true at once. AI returns really are harder to measure than traditional technology returns, and the era of accepting "trust us, it is valuable" is ending. Both facts should shape how you approach the question.
The difficulty is real and well-documented.
of C-suite leaders reported achieving significant ROI from AI, defined as a 20% or greater profitability or cost improvement. Measuring business impact ranked among the most-cited challenges. Forbes AI Study 2025, a survey of 1,075 C-suite leaders.
McKinsey's State of AI research has found that only a minority of organizations can point to enterprise-level earnings impact from their AI deployments. And Deloitte research suggests most organizations need two to four years to realize returns from a typical AI use case, far longer than the seven-to-twelve-month payback that traditional IT investments tend to deliver.
AI often creates value through less tangible channels, better decisions, faster learning, new capabilities, that conventional ROI frameworks were never built to capture.
But difficulty is not an exemption. Boards have watched enough AI spend go out the door that the patience for unmeasured value is gone.
The organizations that will keep funding AI are the ones that can show returns, which means the measurement problem is not academic. It is the difference between a program that survives the next budget cycle and one that does not.
The good news is that a Forward Deployed Engineer, measured correctly, is one of the more measurable AI investments you can make, precisely because it is tied to specific shipped systems.
The three-level FDE value model
The core mistake in AI ROI is measuring the wrong scope. People measure the first system, find the payback underwhelming on its own, and conclude AI does not pay.
They miss that a well-run embedded engagement creates value on three levels, and the shipped system is only one of them. The levels mirror the Find, Ship, Multiply operating model directly, in that order.
-
Level one
Value found
What it measures. The scored roadmap of opportunities the FDE uncovered across the business.
How to quantify it. The expected value of the prioritized roadmap: a reusable asset that de-risks and speeds every future decision.
-
Level two
Value shipped
What it measures. The production system the FDE delivered.
How to quantify it. The specific workflow metric it moved: hours saved, cycle time reduced, error rate lowered, throughput or revenue per person increased.
-
Level three
Value multiplied
What it measures. The acceleration of everything after the first build.
How to quantify it. Time-to-production for initiative two versus initiative one; patterns reused; the team's increased AI delivery velocity.
Measure only the shipped system, and an embedded Forward Deployed Engineer can look expensive against a single use case.
Measure all three levels and the picture inverts, because the roadmap the engagement found and the acceleration it leaves behind are the parts that keep paying after it ends. That is the entire economic argument for buying a capability rather than renting an outcome.
Set the baseline before the engineer arrives
The single most common reason AI ROI arguments go nowhere is that nobody measured the "before."
Six months in, someone asks what changed, and the honest answer is that no one wrote down the starting point, so the debate collapses into competing anecdotes. This is avoidable, and it costs almost nothing to prevent.
Three things to do before the engagement starts
- Pick the workflow metric that the first build is meant to move, and pick one, not ten. If the system is meant to cut claims-processing time, that is the number.
- Snapshot it honestly, including its variability, so you know what normal looks like.
- Agree on the counterfactual: what would have happened without the intervention, so you are not later crediting AI for a trend that was already underway or blaming it for a seasonal dip.
None of this requires sophisticated tooling. It requires the discipline to align technical and business stakeholders on success metrics before, not after, the work.
Research on the organizations that successfully quantify AI value keeps returning to that same point: the discipline around the measurement matters more than the sophistication of the calculation.
An FDE actually helps here, because the Find phase produces exactly the baseline you need. Mapping the workflow and scoring the opportunity forces the "before" number into the open at the start, when it is still easy to capture.
The costs side, honestly
A credible ROI case has an honest denominator. An FDE carries a premium rate; there is no point pretending otherwise. The question is what you are comparing it against, and there are two fair comparisons.
The first is the in-house alternative. Building the same capability by hiring is a real option and a good long-term one, but the current market is punishing: forward deployed engineer is among the hottest roles in AI, with postings up more than 800% in nine months and compensation for the role reaching into the mid-to-high six figures plus equity at the frontier labs.
Matt Paige's analysis of the AI pay gap, drawing on PwC's study of nearly a billion job postings, makes the broader point that the market is paying a steep premium for people who can actually put AI to work. Factor in recruiting time, ramp time, and the risk of a mishire, and the embedded engagement often compares well on a time-to-value basis even at a higher rate.
The second comparison is the one people forget: the cost of another stalled pilot. If roughly 95% of AI pilots deliver no measurable impact, the base rate for an unowned, under-scoped effort is failure, and a failed pilot is not free.
It consumes engineering time, executive attention, and, most expensively, the organization's belief that AI can deliver. Measured against that, the premium for someone accountable to production is often the cheaper path, not the pricier one.
Make the return measurable from day one
Our Anthropic-certified Forward Deployed Engineers deploy Claude into your business and make it stick.
Explore FDE engagementsWhat good looks like
The three-level model is easiest to see in practice. Across our engagements, the strongest outcomes show value on all three levels rather than a single headline number, and reading them that way is the point.
A level-one result is the roadmap: the scored set of next opportunities that makes the following decisions faster and less risky.
A level-two result is a production system moving a concrete workflow metric, the shipped thing doing the measurable job it was built for. Our work with clients including Xometry, Vanco Payment Solutions, and ALTAS AI produced production outcomes of that kind rather than proofs of concept that stopped at the demo.
And a level-three result is the one you feel quarters later, when the second and third initiatives ship faster because the patterns and the team capability were already in place.
Trust underwrites all of it, because a system nobody adopts scores zero on every level. It is worth hearing that in a client's voice: TrueBlue's President and CEO, Taryn Owen, has spoken to the value of that kind of embedded, accountable partnership.
When we cite specific client metrics in a formal setting, we keep them bounded and verified. The pattern, though, is consistent, and it is a three-level pattern, not a single number.
Counting all four kinds of return
One reason AI ROI conversations stall is that people argue about a single hard-dollar number when AI actually produces value in four distinct forms. Naming all four keeps the discussion honest and stops a real return from being dismissed because it did not fit one narrow definition.
Hard financial returns
Direct, bankable gains: cost removed, revenue added, cash freed. This is the number finance most wants, and it usually lives at level two, the shipped system.
Productivity improvements
Time and effort saved across a workflow. Real, but only a return if the saved hours are actually converted into higher output or reduced cost rather than quietly absorbed. Say which, or it does not count.
Avoided risk
Errors prevented, compliance failures averted, decisions made more consistently. Harder to see, often larger than it looks, and especially relevant in regulated work where a prevented failure is worth more than an efficiency gain.
Option value
The capabilities and data foundations that make future initiatives possible. This is almost entirely level three, the multiplied value, and it is the one traditional ROI math ignores entirely even though it is often the largest.
An FDE tends to generate all four, but they land on different levels of the model and on different time horizons.
Presenting them separately, rather than mashing them into one figure, is what lets a skeptical audience see the return that is genuinely there instead of rejecting the whole case because the hard-dollar line alone looks modest in year one.
A measurement routine that survives board scrutiny
Pulling it together, here is the routine that holds up when someone senior pushes on it.
-
Align on metrics first
Get business and technical stakeholders to agree what success means before any building starts.
-
Baseline the before
Snapshot the target metric and its variability, and agree the counterfactual.
-
Measure all three levels
Track found value, shipped value, and multiplied value, not just the system that shipped.
-
Revisit on a cadence
Re-measure at set intervals so the compounding levels actually get counted.
The discipline is the deliverable. A calculation nobody agreed to up front is a debate; a routine everyone signed off on before the work began is evidence.
That distinction is what turns "we think AI helped" into a number a CFO will fund again.
Frequently asked questions
What payback period should we expect from a Forward Deployed Engineer?
The first shipped system should move its target workflow metric within the engagement itself, measured in weeks to a few months. The larger returns, the roadmap that was found and the acceleration that follows, compound over the following quarters. Note that broad research suggests most AI use cases take two to four years to fully pay back, so the embedded model's fast first win is an advantage, not the whole story.
How do we measure soft benefits like team capability?
Through level three: track the time-to-production for your second and third initiatives against the first, count the reusable patterns produced, and note the work your team can now do without outside help. Capability shows up as acceleration, which is measurable even when it feels soft.
What metrics prove an AI project actually worked?
A specific, pre-agreed workflow metric moved in production, not adoption or activity metrics like logins. The organizations that struggle most with AI ROI are usually the ones measuring activity instead of business outcomes.
What if the first use case fails?
A well-run engagement front-loads the risk in the Find phase, so a use case that will not work is more likely to be caught before heavy investment. When something does fail, the found-value roadmap and the baseline discipline mean you learn precisely why, which is itself a return that reduces the cost of the next attempt.
Sources
- Forbes AI Study 2025 (survey of 1,075 C-suite leaders; ROI and measurement findings)
- McKinsey, State of AI 2025 (enterprise-level EBIT impact)
- Deloitte research on AI use-case payback periods (two to four years)
- MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (pilot failure base rate)
- Matt Paige, "The AI Pay Gap," Substack, 2026, citing PwC job-postings analysis
Define the baseline and the metric before any work starts
HatchWorks AI is an Official Anthropic Claude Partner. In a strategy session, our Forward Deployed Engineers will set the measurement up with you, so the ROI is provable from day one.
Book a strategy session


