All articles
Engineering5 min read

Grounding, in practice. How to stop an AI agent from inventing answers

Hallucination is not a model problem you wait out. It is a retrieval and permissions problem you design around. A working checklist for grounding an agent in what your company actually knows.

The first demo always goes well. You point an agent at your website, ask it three questions you know the answers to, and it responds beautifully. The second week is when someone asks about a policy you changed in March, and the agent answers using the version from the page you forgot to delete.

Grounding is the work between those two weeks. It is less about the model than people expect, and much more about what you feed it, what you let it say when the sources are thin, and how quickly stale content leaves the system.

Decide what “the truth” is before you index anything

Most grounding failures are not retrieval failures. They are two documents disagreeing and the agent picking one.

Before connecting a single source, write down which system wins for each category of question:

Question typeSource of truthNot a source
Price, stock, variantsProduct catalog / commerce APIOld landing pages, PDFs
Shipping and returnsCurrent policy pageSupport macros, email threads
Availability, opening hoursCalendar / booking system“About us” page
Technical specsProduct database, spec sheetsBlog posts, press releases
Anything legalApproved legal text onlyEverything else

This table is the actual deliverable. It takes an afternoon and it prevents the majority of confident-wrong answers, because it tells you what not to index. Which matters more than what you do.

Retrieve in units that answer questions

Documents are written for reading top to bottom. Retrieval works on fragments. If you split by character count, you will eventually cut a sentence in half between “we do not offer refunds after” and “30 days, except for defective items”.

Practical rules that survive contact with real content:

  • Split on structure, not length. Headings, list items, table rows, FAQ pairs. A chunk should be able to stand alone as an answer to something.
  • Keep the breadcrumb. Prepend the document title and section heading to each chunk. “Returns → Defective items → EU” is a huge retrieval signal for two extra lines of text.
  • Keep chunks small enough to be specific, large enough to be complete. Roughly a screenful. If a chunk needs its neighbour to make sense, merge them.
  • Store metadata you can filter on: language, product line, region, effective date. Half of “the agent gave the wrong answer” is “the agent gave the right answer for a different country”.

Instruct for refusal, not just for tone

Most system prompts spend their words on personality and almost none on what to do when the retrieved context is insufficient. That ratio should be inverted. The behaviour you want, stated plainly:

Answer only from the provided sources.
If the sources do not contain the answer, say so and offer to connect
the customer to a person. Never estimate prices, dates, availability
or legal obligations. When sources conflict, prefer the one with the
most recent effective date and say that you are doing so.

Then verify it. Ask the agent five questions you know are not covered by your knowledge base. If it answers all five, your instructions are decoration: the model is filling gaps from its general training, which is exactly the failure mode you are trying to eliminate.

Make citations non-optional internally

Customer-facing citations are a product decision; some brands want them visible, some find them clinical. Internally, they are not optional. Every answer should carry the source fragments it used, visible to your team in the conversation log.

The reason is diagnostic. When a bad answer appears, there are only three possible causes, and the citation tells you which one instantly:

  1. The right source existed and was not retrieved → retrieval problem (chunking, embeddings, filters).
  2. The wrong source was retrieved and used → content problem (stale page, duplicate, contradiction).
  3. The right source was retrieved and the answer still departed from it → instruction problem.

Without citations, all three look identical and you end up rewriting prompts to fix an indexing bug.

Treat freshness as a pipeline, not a project

Knowledge decays quietly. The page changed; the index did not. Three habits keep it honest:

  • Re-sync on a schedule that matches how fast each source moves. Catalogs hourly, policy pages daily, static documents weekly.
  • Delete aggressively. An outdated document that is still indexed is worse than a missing one, because it produces an answer instead of an escalation.
  • Keep a small regression set. 30 to 50 real questions with approved answers. Run it after every knowledge change. It takes minutes and catches the “we fixed one thing and broke four” pattern that otherwise only customers notice.

Where grounding ends and permissions begin

Grounding controls what the agent believes. It does not control what the agent does. An agent can be perfectly grounded and still book the wrong slot, because booking is an action and actions need their own boundaries: which tools it may call, which arguments it may fill in on its own, and which steps require an explicit confirmation from the customer or a person on your team.

The two systems are easy to conflate and expensive to confuse. Knowledge answers “is this true?” Permissions answer “am I allowed to?” You need both, and only one of them is solved by better retrieval.


Grounding is unglamorous work: an inventory of sources, a chunking strategy, a refusal instruction, a regression set, a re-sync schedule. None of it demos well. All of it is the difference between an agent your team trusts with customers and one that quietly gets switched off after a month.

Share

Ready to put this
to work?

Bring your use case. We will show the agent handling it live in 30 minutes.

Book a 30‑min demo

Get a call from the agent

Leave your number and the agent calls you back about your enquiry.