The Diagnostic
Anatomy of a confident wrong number by AI - a cash runway question answered with a precise but invented figure.

Why AI gives businesses the wrong number, confidently

By Dancho Dimkov11 min read

A general AI tool will happily answer a business question it cannot actually work out, and hand you a clean, specific, invented number. Here is the simple test that tells you, in advance, which of your questions an AI will get right and which it will quietly get wrong.

A while ago a business owner came to see me, genuinely excited. He had been at a presentation where someone showed off an AI tool, and it had stayed with him. The presenter had typed in an ordinary question, the way you would ask a colleague, the tool answered in two seconds, and it even pointed back to the document the answer came from. He wanted the same thing for his own company, and he wanted it now. "I just want to be able to ask my business anything," he said.

I understood the excitement, because what he saw is genuinely impressive. The technology behind it has a name, RAG, short for retrieval-augmented generation, but you do not need the acronym. All it really means is that the AI is connected to your documents, and when you ask a question it finds the right passage and reads it back to you. For the right kind of question, it is close to magic, and the magic is real.

Then he started listing the questions he actually wanted answered, and I felt the familiar gap open up. Whether one client had quietly grown into too large a share of his revenue. How many months of payroll he could cover if his three biggest customers all paid late. Whether last year's jump in margin was real, or just one unusually good project flattering the average. Almost none of them were the kind of question that demo could answer. And I have seen where that leads often enough to know how it usually ends. Not with "sorry, I could not find that." With a confident, specific, neatly formatted number that was completely made up.

Maybe you have had the same moment. You sit through the demo, the future feels like it has finally landed on your desk, and then you ask it something that actually matters, and the answer is wrong. Not "could not find it" wrong. Confidently, specifically wrong, with no warning at all.

That conversation is why I am writing this. I have watched it happen to enough owners that I no longer treat it as a glitch; it is built into how most of these tools work. For the last stretch I have been building a system for Business Pulse that answers questions straight from a company's own numbers, and the single hardest part was not the clever bit. It was teaching the thing when to say "I cannot answer that."

So let me give you what the demo never does: a way to tell, before you trust it, which of your questions an AI will answer well, and which it will quietly get wrong.

There are two kinds of questions, and underneath they are not the same job

The two kinds of questions: find it (retrieve) versus work it out (calculate), and which side a question is on decides whether AI helps or misleads.

Here is the distinction that makes sense of all of it, and I did not come to it casually. My doctoral research is on how service-based SMEs actually adopt AI and automation, and while I was deep in that work one pattern kept surfacing in how founders talk about it. The questions they most want to put to AI are not all the same kind of question. They fall into two, and which kind you are asking quietly decides whether AI is about to help you or mislead you.

You can ask an AI these two very different kinds of questions about your business. They feel identical when you type them. They are completely different underneath, and the difference is the whole game.

The first kind is a "find it" question. The answer already exists, written down somewhere, word for word. What does our leave policy say about carrying days over to next year. What risks were flagged in the last board pack. What did the client actually ask for in that long email thread back in March. The answer is sitting in a document. The AI just has to locate the right passage and read it back to you. It is genuinely excellent at this. This is the part of the demo that impressed you, and it deserved to.

The second kind is a "work it out" question. The answer exists nowhere. It has to be calculated, from your own messy numbers. Is one client too big a share of our revenue. How many months of cash do we have left at this burn rate. Did our margin actually improve last year, or did one big one-off make it look that way. No document contains that answer. Something has to pull the exact figures out of your spreadsheets, line them up, and do real arithmetic.

Here is the same split, side by side:

The two kinds of questions compared on five criteria: where the answer lives, what the AI has to do, three examples, how good general AI is, and what a wrong answer looks like - find it (retrieve) versus work it out (calculate).

Once you can feel which side of that line a question sits on, most of the mystery disappears.

Why the demo only ever shows the easy half

Demos are built to impress, so they are stocked with "find it" questions. Those look like magic and almost never go wrong. Ask the tool what the contract says about termination, watch it pull the exact clause, and the room nods. Fair enough. It earned the nod.

Now look at the questions that actually run your business. Concentration. Runway. Margin. Trend. Cost per unit. Which product line is quietly losing money. The decisions you lose sleep over are nearly all calculations, not lookups. They are "work it out" questions, almost every one.

So you watch a demo of the easy half and buy the tool expecting it to handle the hard half. That gap, between the half you were shown and the half you need, is exactly where the wrong numbers live. The owner across from me had been shown a flawless "find it" demo, and almost every question keeping him up at night was a "work it out" question. He had not noticed, because nobody had drawn the line for him. It is not that the vendor lied. It is that you were left to assume one tool was equally good at both jobs. It rarely is.

The expensive part is a wrong number that looks right

Here is the part that actually costs money.

When a person does not know something, they tell you. They say "I am not sure, let me check." A general-purpose AI tool does not do that with numbers. It does not have a reliable way to say "I cannot work that out." So it does the only thing it knows how to do. It produces something that looks like an answer: a clean, specific, confident figure that was essentially a guess dressed up as a fact.

A wrong number you can spot is just annoying. A wrong number that looks exactly like a right number is dangerous, because you act on it. You set a price on it. You make a hire on it. You walk into the bank with it. No alarm went off, because the tool sounded just as sure as it does on the days it is right. There is no wobble in its voice, no "roughly," no asterisk. Same confidence, whether the number is real or invented.

That is the quiet risk sitting inside most "ask your business anything" tools. They are brilliant at finding what exists, and dangerous at calculating what does not, and they use the identical confident tone for both.

Why it happens: a language tool doing a numbers job

You do not need to understand the engineering to protect yourself, but a one-line version helps, because it tells you where the edge is.

The tool is, at its core, a language machine. It was built to read and write words. It is extraordinary at it: find the relevant sentence, rephrase it, summarise it, translate it, pull the one clause that matters out of forty pages. When the answer to your question is words that already exist, you are playing to its deepest strength.

A number in a spreadsheet is not language. "Client A billed 41% of revenue last year" is not a sentence waiting to be found and quoted. It is a calculation waiting to be done, correctly, from specific cells. Finding the right words and doing the right arithmetic are two different jobs, and being world-class at the first does not make a tool reliable at the second. When you ask a pure language tool to do maths, it often does an impression of maths. It writes the kind of sentence that a correct answer would look like, with a plausible number slotted in. Sometimes that number is right. Sometimes it is not. The tool cannot tell the difference, and so neither can you.

This is the gap I kept running into in the research and then again in the building. The tools are not stupid and they are not lying. They are being used for the one job they were never built to do reliably, and nobody warned the owner.

What good actually looks like

The fix is not a cleverer guess. It is a tool that knows the difference between the two kinds of question and behaves differently for each.

What good looks like: a find-it answer shows its source, a work-it-out answer shows the figures used. Source or working, never a guess.

For a "find it" question, it finds the passage and shows you the source. Ask it "what delivery time did we promise this client?" and it pulls the exact line out of the signed contract and links you straight to it, so you can confirm in one click.

For a "work it out" question, it does the thing most tools refuse to do. Ask it "is any single client more than a third of our revenue?" and it pulls each client's billings, adds them up, shows you the share and names the client, with the figures it used laid out underneath. And if the numbers it needs were never uploaded, it tells you that plainly, instead of inventing a percentage. Exact, with its working shown. Or an honest "I cannot answer that from your data." Never a confident guess in the space between.

That willingness to say "I do not have that" is the single most important quality a business AI can have, and it is the one almost nobody demos, because admitting a limit does not look impressive in a sales meeting. It is the whole difference between a tool you can hand a real decision and a tool that is right up until the day it quietly is not.

The order matters, and most tools get it backwards

There is a subtler trap hiding underneath all of this, and it is about the order in which a tool decides what to do.

The natural move, for a person and for a simple tool alike, is to reach for the search step first. It is the part that works, and it is the part the demo sold you. You ask a question, the tool goes hunting through your documents, finds something that looks relevant, and answers from it. For a "find it" question, that is exactly right, and it is why the RAG demo was so convincing.

For a "work it out" question, starting with search is precisely how you get a silent wrong number. The search will always find something: a figure in an old report, a number on a slide, a line in last year's accounts. The tool then phrases a confident answer around whatever it found. No actual calculation ever happens. The number was not computed from your real figures, and it was not honestly looked up either. It was improvised from a fragment. And because the search step "worked" in the sense that it found text, nothing anywhere flags that the question needed arithmetic, not retrieval.

So the safe order is the opposite of the intuitive one. A tool you can trust asks itself first: does answering this require a calculation? If it does, it does the arithmetic directly on your actual figures, the same inputs giving the same answer every time, and only falls back to searching your documents when the answer genuinely is written down somewhere. Calculate first. Search second, as the backstop, never the default. A tool that reaches for search no matter what you ask has not solved the hard half. It has just pointed the search step at everything and hoped.

A 30-second test before you trust the next one

You do not need a technical audit to check any of this. You need two questions and half a minute.

The 30-second test: ask a number you already know (expect an exact answer plus source), then ask a number that does not exist (expect I do not have enough data).

First, ask it a number you already know cold. Last year's revenue. Headcount. Your single biggest customer's share. You are not learning the figure; you already have it. You are checking whether the tool gets it exactly right and whether it shows you where it pulled it from. Vague, rounded, or sourceless is a red flag even when it is close.

Then, ask it something that is genuinely not in your data. A figure for a month you never uploaded, a metric you have never tracked. Watch what it does. A tool you can trust will tell you it does not have that. A tool that will eventually cost you will invent something rather than admit the gap. That second answer, the confident invention, is the whole risk in miniature. Better you trigger it on purpose now than discover it the day you are pricing a deal.

What to do next

So before you hand any AI tool a real question, do three things.

Sort the question first. Decide whether you are looking something up or working something out. For "find it," most decent tools will serve you well, and you should use them. For "work it out," raise your guard.

Run the 30-second test on anything you are about to rely on for a decision. One number you know, one number that does not exist. Thirty seconds buys you a lot of certainty.

Ask the vendor the one question that matters: when it does not know, what does it do. If the honest answer is "it answers anyway," you have not found a shortcut. You have found a confident wrong number waiting to happen, and a slower, sourced answer beats a fast invented one every single time a decision is on the line.

None of this means hold off on AI. It means start in the right place. And the right place is not the tool. It is the question.

The mistake almost everyone makes is to start with the technology: see an impressive demo, buy the tool, then go looking for questions to point it at. Turn that around. Write down the handful of questions you actually want answered, the ones that would change a decision. Sort each one into "find it" or "work it out." Only then do you know what you are really shopping for, and whether the tool in the demo can deliver it or never could.

That is not a trick I invented for AI. It is the whole logic of a diagnostic: get the questions right first, and the right tools and answers follow. Get them wrong, and the most advanced technology in the world will just hand you a faster, more confident version of the wrong answer.

If you want the practical map of where to start, we wrote up the first places an SME should put AI to work, and if you want the bigger strategic picture, the three layers of AI value is the piece to read next.

And if you want a system that answers the "work it out" questions properly, from your own numbers, with the working shown, that is exactly the conversation we have in a Business Pulse AI session.

At Business Pulse, that is the one line we will not cross. Every number comes with its source, or it does not come at all.

So which kind of question were you about to ask?

Frequently asked questions