Key points
- An audit starts with a list of tasks and real hours, not a list of tools. A tool without a work inventory is a solution looking for a problem.
- Only the rule-based share of a task is a candidate for automation. Subtract residual oversight from it, and only that difference is a saving.
- The payback calculation is five lines: implementation plus maintenance versus hours times rate. The numbers must be yours, not from someone else’s slide deck.
- Record a baseline before you start: hours, errors, speed, scale, and the date of measurement. Without it you cannot prove the effect to yourself or to your board.
The map from the previous chapter tells you where AI helps. It does not tell you where to start in your business. The lighting store will start with descriptions, because that is where it sinks a week of work into every delivery. A parts wholesaler with six channels will start with feeds. Sequence follows from numbers nobody else holds.
This chapter is a procedure you run yourself: paper or a spreadsheet, a few conversations with the team, zero vendors at this stage. It answers how to implement AI in a company in a way you can account for afterwards: process audit first, tools second. The reverse order ends with buying a solution and hunting for a problem to fit it.
The procedure: five steps you can run on paper
The whole exercise comes down to breaking work into tasks and calculating how much of each you can realistically lift off a person. Go in order. Do not skip step three, because that is what separates an honest calculation from a wishful one.
- List the repeatable tasks of a month, not departments. A task has a start, an end, and a checkable output: "listing a new delivery on the marketplace", not "marketplace operations". Where does the list come from? Calendars, the helpdesk queue, the spreadsheets someone opens every week, and half an hour with each person who actually does the work.
- Assign each task a realistic monthly time. Frequency times the length of one run, plus rework and context switching. If you can, measure for two weeks instead of estimating from memory. Memory understates routine and overstates firefighting, because firefighting is what gets remembered.
- Estimate the rule-based share. The test: how much of this task could a new hire complete given only a written instruction and the data, with no questions about context? Only that share is a candidate for automation. The rest is judgment, and judgment does not evaporate.
- Subtract residual oversight. Someone will review outputs, catch exceptions, and react to changes in the underlying systems. Early on that is substantial. It shrinks only once you have evidence the workflow is stable. Planning for zero oversight is the most common error in these calculations.
- Only the difference is a saving. Record it as a dated assumption with the figures you used, not as a fact. In a quarter you will see how far off you were. That, too, is a legitimate output of the audit.
Step three is the hard one, because it demands honesty about your own process. Instinct says: "we do it the same way every time, so all of it is rule-based". Usually it is not. In a product description, the structure and the attribute fill are rule-based. The decision about what to emphasize on a flagship product is not. We showed what that split looks like across seven concrete tasks, with hours before and after, in a piece on the anatomy of automation savings. Here the principle is enough: you automate part of a task, not a task.
The second criterion: error cost and sales impact
Hours alone will not set your sequence. The task with the biggest hour count can be the worst first choice if its errors are expensive. Compare two slip-ups. A bad description draft gets caught at review and fixed in a minute. A bad price in a feed can sell 200 units below cost before anyone notices.
So give every candidate two more scores, each on a scale of 1 to 3:
- Error cost: does a bad output stay inside the team, or does a customer or a marketplace see it, and can you undo it in a minute or do you untangle it across three systems?
- Sales impact: does the task sit on the path to revenue, for instance by holding up the publication of new stock, limiting channel coverage, or slowing replies to pre-sales questions?
These scores do not decide whether a task is suitable for automation at all. The qualification rule from the previous chapter settled that. They decide the order. High sales impact and cheap errors: front of the queue, even with fewer hours. High error cost: further back. Not because it cannot be done, but because such a task needs more oversight and better data, and both are easier to build once you have one successful implementation behind you.
Three questions that set priority faster than any scoring sheet: does the task return at least weekly, can a bad output be undone in a minute, and can you name the single number that should move after implementation? Three yeses: a first-implementation candidate. One no: the queue. Two nos: it stays with a person for now.
What it means in money: a worked example
Back to the lighting store. The team spends around 55 hours a month on descriptions, variants, and translations. The rule-based share is roughly 40 hours. After subtracting oversight of the drafts, the realistic saving is about 30 hours a month. Now the costs. Implementing the description workflow: EUR 3,000 one-off. Maintenance: EUR 100 a month. All of these numbers are examples, substitute your own.
The calculation: 30 hours times EUR 20 per hour of labor cost gives EUR 600 a month. Minus maintenance, EUR 500 net. The EUR 3,000 implementation pays back after six months. After the first year you are about EUR 3,000 ahead, and the workflow keeps working. This is not a promise of results, it is the skeleton of a calculation. Plug in your own hours from the audit, your own rate, and your own implementation quote, and you will see whether your case adds up, and when.
If you would rather not build the spreadsheet, we made a free Blueprint Check calculator for exactly this. It computes the full cost of a platform together with operational work from your own assumptions, and every amount is editable. A few minutes and you have your own version of the calculation above.
The implementation queue: what goes first, what goes second
The output of the audit is not a ranking but a queue with three tiers:
- First: a task with high repetition, a large rule-based share, cheap errors, and an obvious measure of success. Usually the catalog or feeds, because there you work on data you already hold.
- Second: the higher-impact task that needs integrations, system access, or a data cleanup first. You prepare it in parallel and launch it after the first one, using what the first one taught you.
- Third: the waiting room. High impact combined with high error cost, meaning anything that touches direct communication with customers.
One rule organizes this queue better than all the scores combined: do not launch three implementations at once. Not because of budget, because of attribution. When three things start in the same month and results do not improve, you cannot tell which one failed, and the whole company concludes that "AI does not work here". One implementation at a time: every outcome, good or bad, has a clear cause.
The audit also produces one more thing worth having: a list of tasks you have consciously decided not to automate. That is a real result, because it closes the question and saves you the next round of debate.
The baseline: measure before you start
The most common reason an AI implementation ends without a verdict is mundane: nobody wrote down the state before it started. A quarter later you can no longer reconstruct how long the work used to take, so the discussion about impact collapses into impressions. Record four things for the task going first:
- Hours: what the task takes per month today, from step two of the procedure.
- A quality measure: reworks, channel rejections, complaints in that category.
- A speed measure: time from "work possible" to "work done".
- The owner: who actually does the work today.
Add the context: number of products, orders, channels, and languages, plus the date of measurement. Without it a quarterly comparison means nothing. If the catalog grows by half in the meantime, a drop in hours means something different from what it appears to mean.
Take the follow-up measurement thirty days after launch, not after a week. The first days always look worse: tuning is still in progress, and oversight is at its highest.
Keep the first implementation small
The scope of the first implementation is a strategic decision, not a question of ambition. The recommendation:
- One process, not several at once.
- One owner on your side.
- One number that should move.
- A horizon in weeks, not quarters.
Not "AI transformation of the company", but "cut the time from delivery arriving to being live in two channels". A small scope quickly reveals two things no presentation can predict: whether your data is good enough at all, and whether the team can carry the oversight inside its normal working day.
A large scope takes both of those away from you. Five processes at once: failure cannot be attributed, success cannot be verified, and the company ends up with an opinion instead of a conclusion.
A small start has one more side effect, usually the most valuable. After the first implementation the team can say in its own words what AI does well and what it is not allowed to touch, and the decisions that follow stop being a matter of belief. And the audit almost always ends in the same discovery: the constraint is neither the models nor the budget, but the state of the data they would work on. That is the subject of the next chapter.
Questions
How long does an operations audit take before an AI implementation?
In a small team the task list comes together in one afternoon, and assigning realistic times usually takes about two weeks if you measure rather than estimate from memory. Treat that as orientation, not a norm: with several people and many channels it takes longer. What matters more than speed is that the timings come from the people who actually do the work.
Who should run the process audit?
The people doing the work supply the timings and the exceptions, and one decision-maker owns the priorities and the number that should move. An audit run purely at management level understates routine work, because routine is invisible from above. An audit run purely from the bottom up usually fails to set the sequence, because it lacks the commercial weighting.
What if I have no data on how long tasks take?
Measure for two weeks instead of guessing. A simple log is enough: task, date, start and end time. If measuring is impossible, calculate frequency times the length of one run, add an allowance for rework, and label the result explicitly as a dated assumption. A labelled assumption is useful because it can be corrected later. A number quoted without a source stays in the deck forever.