Just because everyone is racing ahead with AI doesn't mean they're moving in the right direction. The only way to measure ROI is to start with a baseline and a clear objective.
Uber spent its entire 2026 AI budget in the first four months of the year, and its COO admitted publicly that the spending produced no measurable increase in productivity. Amazon has shown the same pattern. Salesforce is projecting $300 million in token spend this year and has yet to show what it bought them. These are companies with finance functions, data teams, and the infrastructure to design and measure AI usage, and they still skipped the basics.
Startups fall into the same trend, but they stay in the background with no press obligation to explain spend, and often no system in place to even notice. In roughly a year of informational interviews and conversations with product managers and engineers at small companies, I haven't met one person who worked somewhere with a way to measure whether AI was moving the needle, or in which direction.
That's not because small companies care less about measurement. Speed beats process by default when a company is small: the cost of skipping a step shows up immediately, while the cost of skipping measurement doesn't show up for months. Uber's own finance team took a whole quarter to catch onto the fact that they burned through their annual AI budget with nothing concrete to show for it. A startup focused on building before planning has the same exposure with none of that infrastructure to catch it at all.
Cheap to build and cheap to productize are two different things. AI made writing code faster, and teams with heavy adoption produced 98% more pull requests while review time rose 91%. That's raw output, and raw output is a small share of what it costs to bring a product to market.
A team that doesn't fully understand code AI generated inherits a maintenance burden it can't estimate, because nobody built a mental model of the system, only a diff. Even a well-maintained build can miss the harder question of whether this was the thing customers wanted to pay for. Ship a feature without product-market fit and the sales cycle lengthens, since sales has to find the narrow segment that wants it instead of selling into demand that already exists, and customer acquisition cost climbs with it. Add enough features fast enough and the product gets harder to use, which shows up as support tickets rather than engineering tickets. Someone has to staff that support function, document the product, and keep the website and app current every time it ships. Raw coding speed is one line item in a much longer list, and it's rarely the most expensive one.
The best place to test the effectiveness of AI is to start with something easy to measure that ties to a clear business objective: bugs caught before release, close rates, tickets resolved without escalation.
A good first metric already gets tracked, produces signal inside one cycle of the work it measures, and would be hard to argue away if it moved. Most of that data exists already, scattered across systems nobody has connected. The baseline is the one piece that can't be recovered later, so it has to be captured before the tool arrives rather than reconstructed from memory afterward.
Running several experiments at once is fine, with one caution: AI added upstream can move a number downstream. If engineering starts shipping faster in the same quarter sales runs a funnel comparison, some of the sales lift may belong to the larger pipeline rather than the funnel. Knowing which experiments touch each other beforehand keeps the read honest.
Most organizations answer "is AI working" by asking people, and people are unreliable narrators of their own productivity. METR found experienced developers were 19% slower using AI while believing they were 20% faster. A 39-point gap between perception and measurement is the argument for comparing two groups instead of surveying one.
Sales is a less obvious place to start, but because commission as an outcome metric already exists, it might be one of the easiest. Deal value and time to close are tracked, comparable across people, and tied to a number nobody argues about. Split the team by willingness rather than by assignment: let the group excited about AI build an AI-assisted funnel, let the group with no interest keep working the way they always have, and compare closed deals over the same period. Match the cohorts on territory quality and seniority so the comparison holds up.
The results may show a lift, or they may not, and either answer is worth having. A lift converts the skeptics without a mandate, because nobody has to trust the technology when they can see a peer's commission check. A flat result, or one trending down, means something in the setup needs optimizing, and the comparison points to where. Without it there's nothing to optimize against.
A negative result is information, not failure, as long as there's a baseline to compare it against.
Check the timeline. If deals take 90 days to close, one quarter is the read. Set that window before starting and hold to it, because judging early produces noise and waiting past it burns runway.
Check the workflow. If more code got written but the added pull requests and review time cancel out the gain, the process around the tool needs work: batch size, who reviews what, which steps still need a human at all. Long reviews can also mean the output quality is poor and reviewers are rewriting instead of reviewing, which points back at the AI output rather than the process. Separate those two before redesigning anything.
Check the success criteria. Success has to be an objective number set at the outset with a baseline behind it: time to first value, customer acquisition cost, support tickets closed without escalation. Without that number, perception fills the gap, and perception is what METR already disproved. Set the criteria, capture the baseline, rerun the test.
Check the use case. The first question with data in hand is what can be optimized, and sometimes the answer is nothing. Starbucks retired an AI inventory system after nine months when baristas spent more time correcting its counts than manual counting would have taken.
Only after those four checks, question the tool itself.
An organization that can prove AI didn't work somewhere is ahead of one that can't prove anything at all. The measurement capability is the asset, and it compounds across every AI decision that follows.
Don't let FOMO get in the way of having a clear strategy. Two weeks of planning costs less than a year of budget with nothing to show for it.