Índice

Índice

A client report marked "Sent to client" with a bar chart, and a magnifying glass showing its +212% figure crossed out in red
A client report marked "Sent to client" with a bar chart, and a magnifying glass showing its +212% figure crossed out in red
A client report marked "Sent to client" with a bar chart, and a magnifying glass showing its +212% figure crossed out in red

The client replies to your recap within ten minutes, asking where "launch moved to the 14th" came from, because nobody agreed to that.

You check. An AI tool wrote the draft, you skimmed it, it read well, but ultimately it was wrong.

It happens. It happens even to the biggest global companies. Deloitte Australia delivered a government report, worth about US$290,000 (A$440,000), that turned out to include references to research papers that don't exist and a made-up quote from a court judgment. Deloitte disclosed it had used Azure OpenAI, corrected the report, and refunded the final payment.

Is AI accurate?

It's accurate at language and unreliable on facts it doesn't have. Grammar, structure, and tone are where modern models rarely slip. Specific facts (a date, a budget, a client's name, who owns what) are different, because a model can't know what it wasn't told, and it won't always say so.

Independent testing backs this up. When journalists from 22 public service broadcasters checked more than 3,000 answers to news questions from four major AI assistants, 45% had at least one significant issue, and 20% had major accuracy problems like hallucinated details or outdated information.

Tools that pull from a set of source documents before answering do better, though not perfectly. Stanford researchers tested leading AI legal research tools, several marketed as close to hallucination-free, and found they still hallucinated on 17% to 33% of queries. GPT-4 on its own did worse, at 43%.

Stacked bar chart of answers from four AI legal research tools. Hallucinated answers ranged from 17% for Lexis+ AI and Ask Practical Law AI to 33% for Westlaw and 43% for GPT-4 without legal sources. Ask Practical Law AI left 63% of answers incomplete.

So is AI accurate? Often. Accurate enough to send to a client without checking? No.

Why a model sounds sure when it's guessing

A wrong answer from AI doesn't look wrong. It arrives in the same calm, confident tone as everything else.

Some of that comes down to how models are trained and scored. OpenAI's researchers compare it to a multiple-choice exam: leave a question blank and you get zero, guess and you might get lucky. Grade models only on how many answers they get right and they learn to guess rather than say "I don't know." Their paper, now peer reviewed and published in Nature, argues this is a big reason hallucinations persist. OpenAI's own summary is the easier read.

Researchers at MIT's CSAIL found the same pattern from a different angle. Standard training for reasoning models rewards right answers and gives nothing for expressing doubt, which makes models "more capable and more overconfident at the same time".

In April 2025, OpenAI rolled back a ChatGPT update that had made it overly flattering and agreeable. A tool that's eager to please you isn't the one you want checking your recap.

Many client work mistakes start with something the AI was never told

The University of Maryland's library guide to AI lists guessing to cover gaps in information as one of the ways it goes wrong. In agency work, that's the big one. A few that will sound familiar:

  • The deadline - The call ended with "let's aim for mid-month." The draft says "launch on the 14th"

  • The budget - Nobody said a number, so the model borrowed a plausible one

  • The reversal - Scope got cut on Tuesday's call. The model only saw Monday's notes

  • The wrong client - Two accounts, similar names, one pasted transcript too many

A mock AI-written client recap email with four mistakes highlighted in red: an invented launch date, a budget nobody agreed, a decision that was reversed on a later call, and another client's name.

"For us it's rarely the model going crazy. It's usually a mistake in the research step. The AI finds accurate information about the wrong thing, like a business or a city with a similar name, or it works from something outdated because the newer version lives somewhere else and it never read it. Then it writes confidently from that, and the draft reads fine."

Jens Rhoades, Owner, Floodlight SEO

In each of these, the model did what it's built to do: fill empty space with something that sounds right. The less it knows about the account, the more space there is to fill, which is why an agency's own context (past calls and decisions) is worth keeping somewhere it can be found.

What changes when the AI has the meeting

If gaps cause many of these errors, the fix is fewer gaps. The Nature paper lists retrieval and tool use among the mitigations that work. In plain terms, give the model the source material instead of asking it to remember or guess.

In client work, most of the source material is conversation. What the client asked for, what you agreed, what changed since. When the AI writing your recap has the notes from that call, plus the project's earlier meetings and docs, it's working from what was said on the call.

That's how Supernormal works. Its meeting notetaker (Mac and Windows) takes notes on your calls in the background, with no bot joining. Supernormal then uses those notes along with your email and docs when you ask it for a status report or a follow-up. If something wasn't settled on the call, like an owner or a date, ask it to flag the gap instead of filling it in. The weekly client status report template is a good place to start, and the guide to client status reports covers what should go in one.

Valerie Dennis Craven, Principal and Strategist at True North Content, gave this a try on a client whitepaper this spring. She put her transcript from the SME interview and other information into Claude and asked for an in-depth outline and brief using only the information she provided. She even asked it to raise questions or leave placeholders for a human to find information. It still got a concept wrong and made its own assumption of how the technology worked. She and the SME didn't need to discuss it, since he, she, and the writer all fundamentally knew it. When she pushed Claude on why it included that, it admitted it "couldn't cite anywhere in my docs that it got the information."

Context cuts errors down. It won't take them to zero. The Stanford legal tools had source documents and still got things wrong. Transcripts mishear names. Someone says "Friday" and means next Friday. A draft grounded in the call is a much safer starting point, and it still needs a person to read it.

Can you trust AI with client work?

For a lot of agencies, the bigger worry is whether they can stand behind what the AI wrote, and explain it when a client asks.

In a Q1 2026 survey of 250 agencies by Digital Applied, hallucination ranked seventh among blockers to rolling out AI, named by 18%. Client trust and explainability ranked second, at 37%. It's research from a company that sells AI services, and the sample leaned toward agencies already interested in AI, so read the exact numbers loosely. The direction makes sense though. Clients want to know where a claim came from.

That's the practical case for drafts built from your own calls. If every line in a recap traces back to something said in a meeting, you can answer "where did this come from?" in seconds.

Five checks before anything goes to a client

"I check AI-assisted work in two ways. A 'stop slop' skill strips out phrasing that reads as AI-generated, and a QA council skill catches any glaring errors… Then I manually review for both tone and accuracy before it ever goes out to a client."

Jenn Feldmann, Senior SEO Strategist, VisualFizz

The checks below are for that last read. They take a few minutes.

  1. Names, dates, and numbers - Check each one against the transcript or doc. This is where plausible-but-wrong hides

  2. Commitments - Every "we'll" and "you'll" should trace back to someone agreeing to it

  3. Anything without a source - If you can't point to where a line came from, cut it or confirm it

  4. The latest version - Make sure a later call didn't undo what this one decided

  5. What should stay internal - Team chat, pricing debates, and opinions about the client don't go in the email

If your team sends client updates every week, write these checks into the SOP for that update so everyone reviews the same things.

The client who caught that wrong date was reading carefully. Start from a draft that knows what was said, and run the five checks before you send. Try the Supernormal meeting notetaker on your next client call.

Perguntas Frequentes

Is AI a reliable source?

Can AI ever be 100% accurate?

Can AI sometimes be wrong?

Can I trust AI to tell the truth?

Does ChatGPT just agree with you?

Junte-se a mais de 700 mil organizações que utilizam o Supernormal

Conclua seu trabalho com clientes num flash com agentes de IA para reuniões e trabalho de projetos.

Comece grátis, sem precisar de cartão de crédito.

Junte-se a mais de 700 mil organizações que utilizam o Supernormal

Conclua seu trabalho com clientes num flash com agentes de IA para reuniões e trabalho de projetos.

Comece grátis, sem precisar de cartão de crédito.