CONTEXT Steering
4 Metrics That Let You Steer an AI Project
Plenty of companies work with AI without being able to say what it delivers. 4 coarse numbers close that gap: usage, saving, quality, cost. What matters is how they are collected. Usage means more than counting logins. The interesting part is how often a conversation with the AI goes back and forth before something is finished. Many rounds are no proof of benefit; they often signal struggle. On saving: 20 minutes times 20 tasks is no saved position. Time gained is scattered time first of all. It counts only once somebody decides what it is used for. There are exactly 2 directions for that: lower cost or higher revenue. Quality shows less in the single mistake than in the mistake that turns up a second time. And cost has to be visible for each work step, otherwise the expensive step stays hidden. Asking these questions before the start is what lets a company steer an AI project by numbers later.
Corporate use of AI has reached the point where the question of return can finally be asked. That is good news: where something gets measured, it can also be negotiated and adjusted. A recent global survey* of roughly 1,700 executives shows where the opportunity sits: 37 percent of organisations report any positive contribution from AI to operating profit. The rest work with the technology while being unable to put a figure on what it contributes. That gap can be closed with 4 numbers any company can gather without a project of its own.
Why is putting a figure on it so hard? The reason rarely lies in the technology. An AI tool delivers its results reliably. What is missing is somebody keeping a record of what it produces and what it costs. In every other area that record-keeping is routine: meter readings for electricity, gas and water get taken and compared. And nobody commissions a project for it. It goes quickly because everyone agreed beforehand on a few meters and reads the same ones every time. An AI project tolerates exactly the same treatment.
4 figures are enough for that, and they are deliberately coarse: usage, saving, quality, cost. A number that can be read off every month in passing is worth more than a metric that needs an analysis of its own. This is our position rather than a standard: so far no framework and no study identifies precisely these 4 figures as sufficient. They come from observing which numbers companies sustain for a year.
Talking about the 4 numbers is one thing, though; collecting them is another. Each figure therefore gets the same 3 steps here: what it claims to measure, where it fails when read naively, and how we collect it in the AIEE programme.
1. Usage
How many people use the thing voluntarily, and how often? This figure is harder to talk up than the other 3, though on one condition. What gets counted has to be actual work: quotations completed, complaints handled, inspection reports produced. Licences issued and logins recorded serve poorly for that. Counting logins is counting parking spaces instead of journeys. The report The GenAI Divide by MIT's Project NANDA from 2025 pins the difference between successful and unsuccessful projects on exactly this point: the ones that delivered measured changes in the workflow, while the others reported the number of licences issued.
A tool used by 3 of the 40 intended people after 8 weeks has a problem. That holds however well it works. One common reason is also the most uncomfortable one: the tool solves a task nobody experiences as pressing. In the surveys a different reason ranks higher, namely the break in the workflow. Anyone who has to switch programs and copy data across by hand for 3 small steps gives up in the third week.
How we collect usage in the AIEE programme
Every participant works on a server environment of their own, and the use of that environment can be counted. 4 figures get looked at:
- Turns per session. A turn is one exchange with the AI: an instruction and the answer to it. The question behind it is how often a conversation goes back and forth until something is finished.
- Requests per week and participant. They show whether the tool sits in the rhythm of the work or gets used in isolated bursts.
- Tokens consumed. A token is the billing unit in which a model measures text, roughly a fragment of a word. The figure says how much text actually ran through the tool.
- Repeat loops. They show how often the same thing gets picked up again before it stands.
Here is the twist in this metric: many turns are no proof of usage. They are frequently a quality problem, because somebody had to follow up 9 times. The same raw number shows either success or struggle, depending on the angle. Which is why every usage figure sits next to the question of whether a usable result came out at the end. Reading both side by side tells you early what the conversation has to be about: the shape of the task, or the way the tool is built into the day.
2. Saving
How much time does one pass save, times how often it occurs? The second half of that question decides everything. A saving of 20 minutes on something that happens 2 times a month stays an anecdote. Worked through as an example, the other case looks like this: a task occurs 20 times a day and saves 20 minutes each time. That comes to 400 minutes a day, or 6 hours and 40 minutes.
And this is exactly where the error of reasoning starts. 400 minutes a day is no position. It is 20 minutes times 20 tasks, spread across several people and across the whole day. Time gained is scattered time first of all. It counts only once it has been bundled and deliberately redirected. Otherwise it evaporates into the working day.
Redirected, it can flow in exactly 2 directions: lower cost or higher revenue. Everything else is a staging post. The real question is therefore less "how much are we saving" than "what do people do with the time gained, and how does that lower cost or raise revenue".
How we collect saving in the AIEE programme
It starts with the baseline, meaning the duration before the launch. It gets measured once, however roughly. Asking afterwards produces the answer that suits the desired result. When moving into a workshop everyone reads the meter. In AI projects that same small step gets skipped routinely.
The saving can be pinned to tasks everybody in a company knows:
- a quotation goes out the same day instead of on the third;
- a case runs through without a stop, because the query back is no longer needed;
- an analysis comes together without input from another department.
Then comes the second half of the exercise, and it is the more uncomfortable one: what was the time gained used for, and who decided that? A saving without an answer to this question stays an arithmetic exercise on paper. A saving with an answer is a decision about cost or revenue.
3. Quality
How often does a human have to correct the output, and how deeply? The rework rate is the share of results somebody still puts a hand to. This is the only one of the 4 that costs ongoing attention. It earns that attention: usage and cost show in hindsight that something has gone wrong. Quality shows it while it is happening.
A rising rework rate at steady usage has several possible causes, and telling them apart pays off. First, the task itself may have changed while the tool lags behind. Second, the model may have changed. Model here means the program trained on data that produces the results. That a trained model loses accuracy once the incoming data shifts is well studied. A paper in Scientific Reports examined 128 combinations of model and dataset from 4 fields in 2022 and found deterioration over time in 91 percent of cases. On top of that, providers swap the version of their model during live operation. Third, it may come down to the users: once past the opening phase they trust the tool with the harder cases.
How we collect quality in the AIEE programme
In our own environment quality arises in the same place as it does in a company: in the work with a tool that produces something. 3 figures get looked at:
- Rework rate. How often a result has to be requested again before it is fit for use.
- Recurring correction. The same mistake a second time. This is the most expensive figure of all, because it shows that the instruction file is failing to bite. The instruction file is the file holding the standing rules for working with the AI.
- Self-catch. How often the AI finds a mistake of its own before a human sees it. This is the only one of the 3 that is allowed to rise.
A mistake that happens once is normal. A mistake that recurs is a design fault. It sits rarely in the model and usually in the instructions.
There is a point here that rarely gets said: every further round of touching a result creates fresh risk of error. Sending something back for a 4th revision buys 4 more opportunities for a new mistake. Past a certain point, reworking stops being quality assurance and becomes its opposite. The value of these figures lies in spotting that point before a customer does.
4. Cost
What does running it cost per month, everything included? The tally holds 4 items: model calls, servers, licences, and the working time spent on upkeep. A model call is each individual request to the program, usually billed by the amount of text processed. That last item is missing from most tallies and gets underestimated routinely. Upkeep here means adjusting templates, chasing errors, working in new cases, supporting users.
As a rule only one half of this is visible. Implementation cost is plain to everyone: consulting, programming, hardware, rented computing power from the cloud. The running cost stays invisible until the invoice arrives. Then it emerges that a considerable sum has gone to AI and cloud providers. And then the attribution is missing: which system, which application, which process, which individual work step. There is an awkward property on top of that: where billing follows usage, the figure grows with success.
How we collect cost in the AIEE programme
That leads to a requirement for how an AI project gets built, and it has 2 parts. A project has to be traceable and it has to be changeable.
- Traceable means the cost is visible for each work step. Knowing it only as a monthly total leaves the expensive step hidden.
- Changeable means the expensive step can be swapped out. There are 3 routes for that: a different model, better handling of the context, or a script in the place where no model was needed at all.
Context here means the amount of text a model reads along with every request. It grows quickly in live use, and it gets paid for again with every call. This is precisely why the cost figure belongs next to usage rather than in a table of its own: only the ratio of the two says whether enthusiasm is turning into economics.
What the 4 Say Together
On their own each figure says little; side by side they say a lot. High usage with rising cost and falling quality is a different picture from low usage with brilliant individual results. The first calls for a decision about money: the shape of the task is right, so the price gets renegotiated or the technology gets swapped. The second calls for a decision about the task.
The obvious objection is that 4 coarse numbers are too thin a basis for a decision about 6-figure sums. There is something to it: a proper business case looks different. It does, however, presuppose that somebody produces it and repeats it. A coarse number available every month for 12 months beats a precise one gathered 2 times. Anyone who later needs the finer calculation already has the series the 4 values provide to build it on.
The benefit lands on both sides of the table. Management gets a basis on which a project can be extended or ended. The department gets numbers to evidence its success. And a provider who supplies these 4 figures unprompted walks into the next negotiation holding the stronger argument.
The 4 Metrics in the AIEE Programme
Pulled together, the collection looks like this:
- Usage: turns per session, requests per week and participant, tokens consumed, repeat loops. Suppose the same analysis takes one participant 40 turns and another 6: the first number points to struggle and the second to practice.
- Saving: duration before the launch, duration after, frequency of the task, decision about the time gained. Example: a quotation goes out the same day instead of on the third and shortens the wait for the customer's reply.
- Quality: rework rate, recurring correction, self-catch. Example: the same formatting error in 3 results in a row is an instruction to the instruction file rather than to the participant.
- Cost: cost per work step, cost per case, share taken by upkeep. Example: a step that reads the same long text along with every call is the first candidate for replacement.
An honest sentence belongs with that: in real projects this kind of measurement is still rarely planned in from the start. It usually reaches the table once the invoice surprises somebody. Rising cost is raising the pressure, though, and that holds for servers in particular. With the pressure, expectations and requirements rise as well.
This is exactly where AIEE comes in. The programme starts less with the question of which model to pick than with the questions asked before the start: what gets counted, what is it compared against, and how is the time gained put to use? Asking those questions early leaves you with steering and transparency later. Asking them late leaves you with an invoice.
* On the state of the sources: the 37 percent comes from the global survey The State of AI in 2026 by the consultancy McKinsey. It polled 1,719 people in 97 countries between 4 May and 8 June 2026. These are self-reported statements by executives rather than audited financial figures. A separate analysis for Germany from September 2026 reports that 43 percent of surveyed companies using AI cannot quantify the contribution to operating profit; that figure appears in the German analysis rather than in the global report. The report "The GenAI Divide" by MIT's Project NANDA from 2025 has no permanent official address and is therefore cited without a link. The paper in Scientific Reports examined classical machine-learning methods on tabular data rather than language models; it evidences the pattern of model ageing rather than its magnitude in today's AI assistants. Reliable independent figures for the share each cost item takes are not yet available. The statistical collection inside the programme follows Section 16 (4) of the participation terms; the content of inputs and the work results are explicitly not evaluated for it. The selection of precisely these 4 metrics is our own judgement: the individual findings on usage measurement, model ageing and billing models are sourced, their combination into exactly 4 figures is not.