“I think we saved a couple hours” isn’t a number. It’s a feeling wearing a number’s clothes.
That’s the trap most people fall into the first time leadership asks about AI ROI. You know the work got easier. You can’t yet say by how much, so you round up to something that sounds reasonable and hope nobody asks a follow-up question. Someone always asks the follow-up question.
There’s no industry standard for how to measure AI ROI at work, as of this writing. I’ve had the direct conversations with AI companies to confirm it. Nobody has agreed on the formula. Which means you get to build it, and you might as well build it in a way that actually holds up.
That’s not just my read. Fewer than a third of companies can measure AI ROI with any real confidence, and the research is clear on why: it’s not that it can’t be measured. It’s that most of them never built the measurement into the rollout in the first place. That’s this whole post in one outside data point.
How to Measure AI ROI at Work: The Formula Is Simpler Than People Expect
A process used to take a certain amount of time manually. AI now automates part of it. Reviewing the output takes a different, usually much smaller, amount of time. The gap between those two numbers is your time saved, and it’s quantifiable the moment you write it down.
That gap is your base ROI. Document it from day one, not after leadership asks. If you wait until you’re asked, you’re reconstructing a number from memory, and memory isn’t evidence.
The Bigger Number Almost Nobody Reports
Time saved gets the attention because it’s easy to explain. There’s a second number that’s usually bigger and almost always skipped: what AI catches that a human process was missing entirely. Is there product sitting on a shelf that never actually made it onto the sales sheet. Are there invoices that have gone uncollected for longer than anyone realized, because nobody had time to cross-check the aging report against what sales thought was already closed. What got flagged once, by someone, and then buried in the shuffle of a normal week, only to resurface because AI actually read the whole file instead of skimming it. That’s where the second number lives, and it’s usually bigger than the time-saved number, precisely because nobody was tracking it before.
That number isn’t about speed. It’s about catching what would have stayed invisible. It deserves its own line, not a footnote under time saved.
But a number needs a comparison to mean anything, and this one doesn’t come with one built in. Time saved has a clear before: the process used to take longer, now it doesn’t. This second number has no before. Nobody was measuring it, so there’s nothing to subtract it from. Which raises the actual question. Bigger than what?
Compared to What
Money that actually came back is the easiest case, and also the strongest. If AI found the uncollected invoice and it got collected, that’s real cash, sitting in an account that didn’t have it before. Count it as what it is. Don’t round it down out of modesty and don’t round it up either. It needs no comparison because the money itself is the proof.
Work your team absorbed without adding headcount is the next easiest, and it’s one you can find without inventing anything. Look at how much one person on your team handled a year ago, compared to how much they handle now, same headcount either way. If you don’t manage a team, look at your own numbers instead. The log you started in the last post already has this in it. Nobody got replaced in either version. The work grew and the team didn’t have to grow with it. That gap, work absorbed against a flat headcount, is a real number, because it’s measured against your own history, not a guess.
Then there’s the number people reach for first and should use least: the pretend employee. “Doing this by hand would have taken three more people,” so AI saved three salaries. That math only holds up if you were actually about to hire those three people, if there was a real job posting or a real budget line with their name on it. If there wasn’t, you were never going to spend that money, so you didn’t save it. Call it what it is instead: something you couldn’t do before and can do now. That’s real. It’s just not a dollar figure, and dressing it up as one is the fastest way to lose the room the moment someone asks a follow-up question.
Work That Never Existed Before
Here’s the case people get stuck on: a process AI built that never had a manual version, so there’s nothing to compare it against. The trick isn’t to check for an old budget line or a backlog ticket. In an AI-first team, the question “can AI do this” gets asked before any paperwork forms, so the paperwork never shows up, even when the demand is completely real.
Check it a different way instead. Name the person who would notice by Friday if this stopped landing in their inbox. If you can name them, that’s real demand, not a guess, and you price it the normal way, against your own log where you can, and as a reasoned, ranged estimate where you can’t. If you can’t name anyone who would notice, that one stays a capability, not a savings number, and that’s fine too.
I had AI build a full PowerPoint deck for a three-session manager training program I actually delivered. I don’t have the design skill to make a deck like that look good on my own, so this is a good test case: some of that time is genuinely mine (I would have built a slower, worse version by hand), some of it belongs to whoever normally cleans up my formatting, and the gap between my best possible effort and what AI actually produced is a capability, not a number I get to claim. One deliverable, split honestly across what’s real and what’s not.
Three tiers, three different kinds of proof. Tier one: money that came back, you count it. Tier two: work absorbed without more headcount, you compare it to your own history. Tier three: capability you didn’t have before, you say exactly that and stop there. Work that never existed before splits across whichever of those tiers it actually earns.
Let AI Ask the Follow-Up Question First
You already know AI will agree with you if you let it. That was the whole point of the very first thing you fixed about it, back when you stopped letting it be a cheerleader. The same discipline applies here, and it matters more here than almost anywhere else, because this is the number someone is going to challenge.
Before you bring a savings number to leadership, hand it to AI first and ask it to argue with you. Which tier is this actually in. Is there a real number behind it or a guess dressed up as one. If you’re claiming a hire you avoided, ask it to check that against the job you were actually about to post, not the job you can now imagine skipping. Point it at the log you’ve been keeping since you started tracking your own time, and have it build the comparison from what’s actually in there, not from what would sound impressive in the meeting.
This isn’t a new skill. It’s the same audit-file habit from earlier in this series, aimed at your own numbers instead of someone else’s build. The smart part was never doing the math. Excel already does the math. The smart part is knowing which comparison is honest, and that’s exactly the question AI should be answering before leadership asks it.
It’s Not Just Your Time
If you build a workflow and only you use it, the ROI is your ROI. If your team uses it too, the time saved multiplies across every person running it. Track adoption the same way you track hours. A workflow that saves thirty minutes a day for one person is a nice result. The same workflow used by a whole team is a business case.
For Engineering and Production Teams, Track These Too
Time saved isn’t the whole picture if your team ships code or produces creative work. There’s a specific set of numbers worth tracking there, and none of them are complicated once you name them.
Time to PR is the first one. A pull request, PR for short, is what a developer submits when a piece of code is ready for someone else to review before it gets folded into the shared project. Time to PR is how long it takes from starting a task to having that request ready. If AI is helping write and review code as it’s drafted, this number should shrink.
Time to merge comes next: how long a PR sits before it actually gets approved and added into the shared project. This is often more about review bandwidth than coding speed, but AI can speed up the review itself too, summarizing what changed, flagging risk, catching what a tired reviewer might miss.
Time to release: how long from merged code to something a user can actually touch. This is the number leadership tends to care about most, because it’s the one tied to when value actually reaches a customer.
Bugs over time matters just as much as speed. Faster isn’t better if it’s shipping more broken, and this isn’t hypothetical: a 2026 benchmark report found that while AI-assisted throughput often looks great on paper, incidents per pull request rose over 20 percent and change failure rates climbed roughly 30 percent at the same companies, because review and quality checks didn’t scale with the extra output. If AI is catching issues a human process missed, same argument as “The Bigger Number Almost Nobody Reports” above, that should show up here as fewer bugs making it to release, not just raw speed. Track it either way, because right now the industry data says most teams aren’t.
Then there’s the harder number: how much you’re actually delivering, weighted for difficulty. Counting completed tickets or requests on their own is misleading, because a five-minute fix and a two-week rebuild both count as “one thing done” if you’re only counting. Weight each use case by the effort it actually took, a rough tier system works fine: light, medium, heavy, and compare weighted output now against weighted output before AI entered the picture. That’s the real “are we producing more” number, not a raw count that can be gamed by knocking out the easy tickets first. It’s also the one number in this section a production or creative team can use directly, the same weighting works whether the unit is a code change or a finished asset.
Most of this lives in tier two from “Compared to What”: measured against your own team’s history, not a guess about headcount you never actually had the budget for. The bug number is the exception. It’s the engineering version of the second number back in “The Bigger Number Almost Nobody Reports,” and like that one, it earns its own line instead of a tier.
What Happens When Leadership Doubts the Number
Here’s what actually happens, in my experience, not the version people expect. Leadership doesn’t usually say the number is impossible. The skepticism is more specific than that. Is the number real. And is there time that can actually be redirected out of it, or does the saved time just evaporate.
When my own nine-hours number, the one you’ll see broken down in a moment, got that exact reception in a meeting, I said nothing back to the doubt. Arguing about whether a number is real is a conversation that goes nowhere. I kept logging, and I let the logs stand for themselves. The number you can prove is worth more than the number you can claim, and proving it takes patience, not a better argument.
Where the Time Actually Goes
The honest answer to “does the time go anywhere” matters, because a lot of people can’t answer it. Mine doesn’t disappear into free time. It goes into running training for other people, thinking through what comes next, and directing AI to execute on it. Not passive time. Redirected time. Most people don’t free up nine hours and then go looking for something to fill them. They take on the next thing before they’ve even registered the time as extra, which is exactly why the number keeps climbing instead of flattening out once the obvious backlog clears.
I have AI log all of it now. Every PowerPoint, every Excel file, every dashboard or report I send out. Workflows like manager feedback loops and SharePoint builds. How many hours each one saved when AI produced the output instead of me building it by hand.
That log is where my own proof point lives. I used to call it tier two only: my own output measured against my own history, and nothing more, so nobody could accuse me of pretend-employee math. I’ve come around on part of that. Tier two was never just a personal productivity stat. Look at what it actually measures: the work grew and the team didn’t have to grow with it. If I’m consistently producing another full person’s worth of output at this desk, that’s a potential additional hire not brought in. The difference is I’m not inventing the headcount. I’m counting it from the log. And I’m still not claiming a salary saved, because there’s no job posting and no budget line, and the rule from earlier still applies. What I’m claiming is the capacity, with the receipts behind it.
Nine additional hours of output per day. Same person, same desk, dramatically more coming out of it: decks, trackers, dashboards that would have taken weeks to build without AI.
Here’s one real day out of that log, not the average, just an example of how the hours actually add up: a training deck that used to take three hours took twenty minutes of review. A tracker update that used to take an hour took ten. A status report, ninety minutes down to fifteen. None of those numbers are dramatic on their own. Stacked across a normal day of them, they’re where the nine hours comes from. That’s the whole point of a log: it’s not one big number you assert, it’s a lot of small ones you can each stand behind.
Leadership doesn’t fully buy that number yet. I understand why. It sounds made up until you’ve watched it happen to your own work. I’m not trying to win the argument in one conversation. I’m letting the log make the case slowly, one entry at a time, because that’s the only version of this argument that actually holds up.
Next up: you’ve proved your value. The next question is where this goes from here, and how to make sure you’re building for more than just yourself.




Leave a Reply