How do you measure AI ROI in a small business?
Short answer
Pick one process and record what it costs: the hours, the errors, the time each job takes and the cost per order. Change it, then count the same things for the same length of time. The difference is the return. It counts as money only where a cost stopped, a planned hire was skipped, or more work got done.
The question sounds like it needs a finance team. It does not. It needs one process and a few numbers you can collect yourself. Then it needs the discipline to count only what changed.
Start with the process, not the tool. "We bought an AI assistant for everyone" is hard to measure. Nobody wrote down what the old week looked like. "Wholesale orders stop being retyped into the ERP" is easy to measure. You can count the orders, the minutes each one takes and the mistakes that follow.
A company of 10 to 200 people has an advantage here. The person who runs the process is often one desk away. You can ask how long a step takes, then watch it for a week to check.
Not sure which process to start with? That is what the AI audit call is for. We name the two or three use cases we would start with. We also say where a simpler fix would beat AI.
What should you record before the project starts?
Short answer
Record how the process runs now, before anything changes. Note the hours it takes each week, who does it, and what the tools cost. Note how often it runs, how often it goes wrong, and how long it takes from start to finish. Collect it over a few normal weeks, so one odd week does not set the number.
This record is the baseline, and it is the step people skip. Without it, the only proof afterward is how the team feels. Feelings about a new system can run hot in the first month and cool by the third.
Keep the list short:
- Hours per week, by person
- How many times the work runs
- Errors per hundred, and the cost of each
- Time from request to done
- Software and outside service costs
Write down where each number came from. A count from the order system beats a guess from memory. Anything you cannot get stays marked unknown. An industry average in the gap makes the result look precise when it is a guess.
For ecommerce work we publish our full method. It runs in six steps, from mapping the current process to comparing the new one against the original baseline. You can read it on the ecommerce operations page, word for word as we use it.
If an outside firm does the work, read what an AI consultant does for a small business before you hire. To get the team using the result, see AI training for employees.
How do you calculate the return on an AI project?
Short answer
Work out the monthly cost of the process before and after, with the same measures for both. Subtract to get the monthly saving. Then compare it with what the change cost to build and run. Saved hours count as money only when they become a cost that stopped, a hire you skipped, or work you can now take on.
Here is an illustrative example, not a client result. An order desk retypes 600 wholesale orders a month from email into the ERP. Each one takes about six minutes, so that is 60 hours a month. At $35 an hour, the step costs $2,100 a month. Three orders in every hundred go out wrong. Each one costs about $40 to put right, which adds $720.
Now the email and the ERP are connected. A person handles only the odd cases, about 10 hours a month. Wrong orders drop to one in a hundred. The step now costs $350 in labor and $240 in fixes. On paper, that is $2,230 a month saved.
The words "on paper" matter. Hours saved by people who stay on the payroll are not cash yet. Something else has to change, and that change is what you count:
- Count the hire the desk no longer needs.
- Count the margin on extra quotes the team sends.
- Count zero if the hours vanish into slower afternoons.
In the last case, only the error savings are left. Count one of the three for each process, never two. That way the same hours never get counted twice.
What counts as a return for a 10 to 200 person company?
Short answer
Four things count. The first is hours your team gets back. The second is fewer errors and less spent fixing them. The third is a shorter wait from request to done. The fourth is a lower cost per order or per case. Time back is what owners feel first, and it becomes money when it replaces a hire.
Time back is the easiest to see and the hardest to bank. Broad research shows how small it can look on average. In a Federal Reserve Bank of St. Louis survey of U.S. workers, those who used generative AI saved 5.4% of their work hours. On a 40 hour week, that is about 2.2 hours.
Source: Federal Reserve Bank of St. Louis, The Impact of Generative AI on Work Productivity, 2025.
One process rebuilt end to end can return far more than two hours a week, or nothing. That is why the measure belongs to the process, not the company.
Errors can be where the cash shows up first. A wrong shipment means postage twice, a credit and a phone call. The wait from request to done matters when customers are the ones waiting, for a quote, an invoice or a refund. Cost per order ties the rest together. It still means something when volume grows.
How long until you know whether it paid off?
Short answer
Plan on two periods of the same length. One runs before the change, and one runs after it has settled. A process that runs every day can give a clear answer within a few months of going live. Work that peaks at month end or in one season needs a full cycle on each side before the comparison is fair.
The first few weeks after launch are not a fair test. The team is still learning the new steps. The cases nobody planned for are still turning up. Let the change settle, then start the clock on the after period.
Match the periods to the work. If the before period was a quiet month, the after period should not be your busiest one. If the business changes in the middle, write it down beside the numbers. A new customer or a price change can move the result on its own.
Keep checking after the first result. A system that saved hours in its first quarter can drift as the business changes. The record you took at the start is what shows the drift.
What if the numbers say it did not pay?
Short answer
Then say so, and decide what to change. The time may have gone into other work instead of off the payroll. The process may have been the wrong one. A step may need removing, not automating. A plain no on one use case is a useful answer. It keeps the budget for the use case that will pay.
Flat results are common enough to plan for. A National Bureau of Economic Research study looked at workers in Denmark who used AI chatbots. It found no measurable change in earnings or recorded hours, and ruled out effects larger than 2%. In the full paper, most chatbot users said they moved the time they saved to other tasks.
Source: National Bureau of Economic Research, Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI, 2025.
That is the trap the counting rule is there to catch. Time saved and moved to other tasks leaves no mark on the books. The tool did not fail. The project never named where the hours would go.
When a result comes back flat, check three things before you drop the idea:
- Was the before record good enough to compare?
- Did the saved hours have somewhere to go?
- Was this the costliest process, or just the easiest?
Sometimes AI was not the fix. The better change can be deleting a step, changing who does what, or switching on a tool you already pay for. When AI is the right tool, we build it. Then we check it against the record taken before we started.


