on the record · Sep 2026

Confidence is the new bottleneck

Ten years in enterprise AI pre-sales — the technical chair, not the selling one — and what replaced the education problem. Literacy got solved. Permission inverted. Calibration got worse, and that is now the expensive one.

Sep 202614 min readBy gopal joshi, Founder, Stringify AI

About ten years ago I sat in a conference room with a whiteboard and a sponsor who controlled a budget. I was not the one selling. I was the technical half of the room — brought in to make the case credible, and then to live with whatever got promised.

That turns out to be the better seat for noticing what goes wrong. The person asking for the money moves on to the next opportunity; the technical lead inherits the commitment and finds out, twelve months later, exactly which sentence in that room was too confident.

What follows is not one account. It is the pattern across a decade of those rooms — different companies, different sponsors, the same four or five failures, watched from beside the person doing the asking. I have only recently started doing the asking myself, which is what prompted me to write it down.

I spent the first meeting explaining what a model was. Not the architecture — the idea. That you could take a decade of operational history and produce a number that was useful before the thing it described had happened. The second meeting was about why that number would be better than the forecast his team already produced in a spreadsheet every Thursday. The third was about what happens when it's wrong.

Only in the fourth meeting did we talk about money.

That was the job. Most of the cycle wasn't selling — it was building the vocabulary the sponsor needed in order to evaluate what we were proposing. That work fell to the technical chair, because it was the only chair in the room that could answer the follow-up question. And underneath the vocabulary was something he never said out loud: he was being asked to take a career risk on a category he couldn't independently assess. If it worked, the credit would be diffuse. If it failed, the memory of who signed would be very specific.

So I taught. Slowly, and from the seat beside the one asking for the budget. It worked often enough that teaching became the shape of the job.

I want to describe what has changed since then, because the popular version of the story is wrong in a way that costs money — including, right now, probably yours.

The popular version goes: consumer AI solved the education problem. Every executive has now used a chat assistant. The thing I spent three meetings on is ambient. You can walk into a room and start at the proposal.

That's true. It is also the least interesting true thing about the last few years.

Three separate things were happening in that old conference room, and because they used to move together, people now assume they still do. They don't. They've come apart, and they're moving in different directions.

Literacy — does the sponsor have the concepts to follow the argument? This is solved. Completely solved, and not by anything the enterprise industry did. It was solved by hundreds of millions of people getting a free tutor and using it.

Permission — is the sponsor allowed to say yes without risking their standing? This has not just improved, it has inverted. The defensive posture used to belong to the person who wanted to spend money on AI. It now belongs to the person who doesn't. "We're evaluating" is a phrase that buys you a quarter, not a year. This is the real unlock of the last three years, and it is worth naming precisely, because it is what people actually mean when they say the education problem is solved. They're not describing knowledge. They're describing cover.

Calibration — does the sponsor's mental model of the technology match what the technology will actually do inside their organization?

That third one got worse. Measurably, expensively worse. And it got worse because the first two got better.

Why a solved education problem creates a harder one#

Here is the mechanism.

A sponsor in 2016 had a vacuum where their model of AI should be. Vacuums are easy to work with. You can fill a vacuum. The person knows it's empty, they know you know more than they do, and the entire interaction has an agreed shape: you explain, they ask, you explain again.

A sponsor today has a fully formed, experience-backed, high-confidence mental model of AI. They didn't get it from a vendor deck. They got it from six months of daily use, which is the most persuasive teacher there is, because it isn't a claim — it's a memory of something that worked.

The problem is that the experience they had is a poor guide to the thing they are about to buy. Consumer AI and enterprise AI share a component and almost nothing else. The sponsor has learned, correctly and deeply, a set of lessons that are wrong at their own organization's scale.

And you cannot fill a space that is already full. You have to displace something, and the something has evidence behind it. That is a fundamentally harder conversation than teaching, and it is harder in an interpersonal way, not just an intellectual one — which I'll come back to.

The five beliefs consumer AI installs#

These are the specific ones. I see them in nearly every room now, in roughly this order of frequency. If you are a sponsor, read these as a diagnostic on yourself.

1. "The demo is the product."#

Where it comes from: In consumer AI there is no gap between demo and deployment. You type, it answers, and that is the product. There is no integration surface, no permissions model, no one else's data, no second user with different access rights. The distance between "I saw it work" and "it works" is zero.

What's actually true: In an enterprise deployment, the demo is a modest fraction of the work — often something like a fifth of it, sometimes much less. The rest is: getting the system access to the data it needs, in a form it can use, that is fresh enough to be true. Deciding who is allowed to see which outputs. Building a way to measure whether it's right. Handling the cases where it isn't. Getting a security review through. Then the part nobody budgets for, which is convincing forty people to change how they do a thing they already know how to do.

What the belief costs you: Timelines quoted at demo speed. A pilot that goes brilliantly in six weeks, followed by a production effort that takes nine months, followed by a meeting in month four of those nine where a senior person asks why this is taking so long, and the honest answer — "because the demo was never the hard part" — sounds like an excuse rather than an explanation. Trust drains out of the project right when it needs the most.

2. "Prompting is the skill."#

Where it comes from: In consumer use it really is the skill. Better instructions, better answers. That's a real, learnable, satisfying loop and people are right to have learned it.

What's actually true: At enterprise scale the leverage moves almost entirely to context. Not how you ask, but what the system can see when you ask: which documents, how current, which of the four conflicting versions of the policy is the authoritative one, whether last quarter's restated numbers made it in. When an enterprise system gives a bad answer, the cause is usually not the phrasing of the request. It's that the system retrieved the wrong thing, or the right thing from eleven months ago, or nothing at all and answered anyway.

What the belief costs you: Organizations staff for the visible skill and skip the invisible one. They hire or train people to write prompts and underinvest in the data work that determines whether any prompt could have worked. Then quality plateaus somewhere unsatisfying and nobody can explain why, because the team is optimizing the wrong variable with great discipline.

3. "It was right when I used it, so it's reliable."#

This is the expensive one.

Where it comes from: Personal use has an error-correction loop so tight it's invisible. You read every single output. You are, usually, the domain expert on what you asked. And the cost of a wrong answer is that you notice, sigh, and retype. Under those conditions a system that is wrong one time in twenty feels excellent — because you caught all five percent, instantly, for free.

What's actually true: Every part of that loop breaks at scale. Nobody reads every output. The person receiving the output often isn't the expert who could spot the error. And a wrong answer doesn't stop — it flows downstream into a report, a decision, a customer conversation, a filing.

Do the arithmetic once and it changes how you buy. Take a process that runs 10,000 times a month with a 5% error rate. That's 500 wrong outputs a month. If each one takes twenty minutes to detect and correct, you've created roughly 167 hours of cleanup — a full-time person, hired to follow your automation around. And that's the good case, the one where you catch them.

So the question that matters is not "how accurate is it." It's: what happens to a wrong one? Does it get caught? By whom? Before or after it reaches someone who acts on it? A 90%-accurate system with a real review path can be an excellent investment. A 98%-accurate system whose errors go straight into production undetected can be a liability, and the second one demos better.

What the belief costs you: Nobody funds evaluation, because evaluation feels like paying to be told something you already believe. Then a wrong answer surfaces publicly, and the organizational response is not proportionate — it's a hard stop. I've watched a single visible failure end a program that was, on the numbers, working. Confidence that was never grounded in measurement has nothing to fall back on when it's punctured.

4. "It'll be cheap. My subscription is twenty dollars."#

Where it comes from: Consumer pricing is flat, subsidized, and — this is the important part — it hides everything that isn't the model. You never see the cost of the infrastructure, the safety work, or the evaluation, because you're not paying for those. You're paying for a seat.

What's actually true: In an enterprise deployment the model is frequently one of the smaller line items. The budget is dominated by the surrounding system: data engineering, integration, evaluation and monitoring, human review capacity, security and compliance review, and then the permanent, unglamorous cost of keeping it working while everything underneath it changes.

What the belief costs you: Budgets sized for a tool when what you're buying is an operating model. This produces a specific and common failure: a well-funded build followed by an unfunded year two, in which the thing quietly degrades because no one owns it, and then gets switched off and remembered as a failure of the technology.

5. "The models are commoditized, so we'll build it ourselves."#

Where it comes from: This one's grounded in a correct observation. Capable models are broadly available. The layer everyone was fighting over five years ago really has flattened.

What's actually true: The observation is right and the conclusion doesn't follow. The model being a commodity is precisely why it isn't where the difficulty lives. What's hard is your data, your workflow, your definition of an acceptable answer, and your appetite for a wrong one — none of which are commoditized, and all of which you'd have to solve whether you built or bought.

The build/buy question also changed shape. The dominant cost is no longer construction, it's maintenance: the model underneath you gets deprecated, the retrieval quality drifts as your documents change, and your evaluation set goes stale the moment your business does. "Can we build this?" is usually yes. "Can we still be running this in three years?" is the question that decides it.

What the belief costs you: Eighteen months and a small team, arriving at something roughly equivalent to a v1 you could have bought, with no evaluation harness and no one left who remembers why the retrieval was configured that way.

Correcting confidence is a status transaction#

Here's the part that isn't about technology.

In 2016 I was a teacher in that room, and teaching is a socially comfortable role — especially from the technical chair, which carries no ask. The sponsor had agreed to be taught, implicitly, by taking the meeting. Nobody lost anything.

Today, and now from the chair that also carries the ask, explaining why the thing a sponsor confidently described won't work the way they think is not teaching. It is correcting — often in front of their team, about a subject they have recently been telling people they understand. That is a status transaction, and it goes wrong in two directions.

Do it bluntly and you lose the room. Not the argument, the room. The sponsor stops being your collaborator and starts being someone defending a position, and every subsequent piece of evidence you present gets read as an attack.

Avoid it and you win the deal on terms that cannot be met. Which is worse — and I have watched that bill arrive, from the seat that had to pay it. The sponsor's confidence becomes the delivery team's commitment, twelve months out, and by then nobody remembers who was optimistic in the room.

The move that actually works is to stop arguing and start instrumenting.

You cannot talk someone out of a belief they formed through experience. You can only give them a better experience. So: take the smallest real version of the thing. Not a demo — a real workflow, real data, real edge cases, run for a bounded period. Count what it gets wrong. Put the number on the table without commentary.

The number does the work that no amount of explanation will. And crucially, it does the work without anyone having to be wrong in front of their team, because the sponsor gets to arrive at the conclusion themselves, which is the only way conclusions of this kind ever stick.

This is also why I've become suspicious of my own eloquence about it. If I find myself explaining at length why a sponsor's confidence is misplaced, I've usually chosen the harder path over the shorter one. The shorter one is six weeks and a spreadsheet.

What to ask now, if you're the one signing#

The old buying questions are dead. "Can AI do this?" has an answer now, and the answer is almost always a qualified yes, which means the question no longer discriminates between good and bad ideas. It selects for nothing.

Five questions that do:

  • What does a wrong answer cost, and who catches it? If you can't name the person or the mechanism, you don't have a system, you have a demo with a budget.
  • What's the measured accuracy on our data — not the vendor's? Demo data is curated by definition, often unconsciously. The number that matters comes from your own messy, contradictory, half-migrated reality.
  • What does this need to see, and are we permitted to let it see that? Most enterprise AI projects that die quietly die here, in month three, in a room with the security team. Find that wall in week one.
  • Who owns this in month nine? When the underlying model is deprecated, when the document set has drifted, when the person who built it has moved teams. If the answer is vague, price in a rebuild.
  • What's the smallest version that produces a real number in six weeks? If there isn't one, that's information. It usually means the scope is a category rather than a use case.

Notice that none of these are questions about AI. They're questions about operations, ownership, and error tolerance. That's not an accident. The technical question stopped being the hard question, and the buying process at most organizations hasn't caught up.

The scarce thing changed#

Ten years ago, the scarce resource in that conference room was imagination. My job was to expand the sponsor's sense of what could be done, because everything was bounded by what they'd seen before and they'd seen almost nothing.

The scarce resource now is discrimination. Not imagination — everyone has that; there are eleven proposed use cases on the whiteboard before I sit down. What's scarce is the ability to tell a real one from a plausible one. To look at eleven and say: two of these will work, three will work but aren't worth the change management, five are demos that will never survive contact with your actual data, and one is an actively bad idea that will make things worse.

So the most valuable thing I can do for a sponsor now is subtract. Narrow the surface. Kill the seven that shouldn't exist and put the whole budget behind the two that should.

That's an uncomfortable pitch. It's shorter than the old one, it sounds less ambitious, and it asks the sponsor to spend less than they arrived prepared to spend. It's also, by a wide margin, the highest-return conversation available in the room.

For ten years I spent the first three meetings explaining what was possible, while someone beside me asked for the money. I now ask for it myself. And I spend those meetings explaining what isn't.

Why this is the company I'm building#

Four of those five buying questions are the same question wearing different clothes: can you show me what it actually did?

What does a wrong answer cost and who catches it — that's a record. What's the measured accuracy on our data — that's a record. What is it permitted to see — that's a record. What's the smallest version that produces a real number in six weeks — that is a record you can get before you commit.

That is the whole of what an Inspector does. It sits in front of the AI already running across your systems and keeps the record: what went in, what was filtered before it reached the model, what came back, what left. It starts in Watch mode — read-only, touching nothing, on your own infrastructure — and reports what it would have caught.

Which is the six weeks and a spreadsheet, built as a product. Not an argument that your confidence is misplaced. A number, on your own data, that you arrive at yourself.

If it turns out your calibration was right all along, you've lost six weeks and gained an evaluation harness you were going to need anyway. If it wasn't, you found out at the cheapest moment it was ever going to be available to you.

Either way you stop guessing, which is the only outcome I'd defend.

← On the record

On the record

For ten years I explained what was possible, while someone beside me asked for the money. Now I ask for it myself — and I spend those meetings explaining what isn't.
gopal joshi
Founder, Stringify AI
See what an Inspector catches in Watch mode →
gopal joshi, Founder of Stringify AI