AI

How to Build an AI Chatbot That Answers From Your Own Documents

October 2, 20267 min readBy Saad Minhas

RAG · AI Chatbot · LLM · pgvector · Embeddings · Knowledge Base

How to Build an AI Chatbot That Answers From Your Own Documents

"Can we train ChatGPT on our data?" is one of the most common questions I get from clients. It's a fair question, and the answer is usually: you don't need to train anything.


What you want is a chatbot that reads your help docs, policies, product manuals or internal wiki, and answers questions using only that information, with links to where it found the answer. The technique for that is called RAG, retrieval-augmented generation. The name is worse than the idea.


This guide explains how it works in plain terms, shows the code for each step, and covers the parts that take most of the time in a real project, which are not the parts most tutorials show.


Why not just train a model?


Fine-tuning changes how a model writes. It's good at teaching a style or a format. It's bad at teaching facts, because the model can still mix up or invent details, and every time your documents change you'd have to train again.


RAG keeps your documents separate from the model. When someone asks a question, you find the most relevant passages and hand them to the model along with the question. The model's job is just to read those passages and answer. Change a document, and the next answer uses the new version.


How RAG works, in five steps


  1. Split your documents into small chunks.
  2. Embed each chunk, which turns its meaning into a list of numbers.
  3. Store those numbers in a database that can search by similarity.
  4. Search for the chunks closest to the user's question.
  5. Answer by sending the question and those chunks to the model.

Steps 1 to 3 happen once per document (and again when it changes). Steps 4 and 5 happen on every question.


Step 1: Clean and split your documents


This is the step that decides whether your chatbot is good or not, and it's the one people rush.


Before splitting, clean the text. Remove navigation menus, cookie banners, repeated headers and footers, and page numbers from PDFs. If a human would skip it while reading, the chatbot should never see it.


Then split by structure, not by character count. Headings are natural boundaries. A chunk that holds one full section ("Refund policy for annual plans") is far more useful than a chunk that starts halfway through one paragraph and ends halfway through the next.


Some rules of thumb that work well:


  • Aim for chunks of a few hundred words. Big enough to hold a complete answer, small enough to stay on one topic.
  • Overlap neighbouring chunks by a sentence or two so nothing important gets cut in half.
  • Keep metadata with every chunk: document title, section heading, URL, last updated date, and who is allowed to see it.

That last field matters more than anything else in this guide. We'll come back to it.


Step 2 and 3: Embed and store


An embedding model turns text into a vector, a long list of numbers. Texts with similar meaning end up with similar vectors, even if they use different words. "How do I get my money back?" and "refund policy" land close together.


If you already use Postgres, you don't need a new database. The pgvector extension adds vector search to it:


create extension if not exists vector;

create table chunks (
  id bigserial primary key,
  document_id text not null,
  title text,
  url text,
  team_id text not null,
  content text not null,
  embedding vector(1536)
);

create index on chunks using hnsw (embedding vector_cosine_ops);

Creating the embeddings is a single API call per chunk. Here it is with OpenAI's small embedding model, which outputs 1,536 numbers per chunk:


import OpenAI from "openai";

const openai = new OpenAI();

async function embed(text: string) {
  const res = await openai.embeddings.create({
    model: "text-embedding-3-small",
    input: text,
  });
  return res.data[0].embedding;
}

Embeddings are cheap compared to the chat model calls you'll make later, so don't over-optimise this step. Embed everything, and re-embed a document whenever it changes.


Step 4: Search for relevant chunks


When a question comes in, embed it with the same model and ask the database for the closest chunks. In pgvector, <=> means cosine distance, so smaller is closer:


select title, url, content
from chunks
where team_id = $2
order by embedding <=> $1
limit 5;

Notice the where team_id = $2. That's the permission filter, and it has to live in the query itself. If you search everything and filter afterwards, or worse, ask the model to "ignore documents the user can't see", you will eventually leak something.


Vector search has one known weakness: exact terms. Product codes, error numbers, people's names and SKUs often don't have a strong "meaning", so similarity search can miss them. The common fix is hybrid search: run a normal keyword search alongside the vector search and merge the results. Postgres can do both.


Step 5: Answer from the context


Now send the question and the retrieved chunks to the model, with clear instructions about how to behave. Here's a version using Anthropic's SDK:


import Anthropic from "@anthropic-ai/sdk";

const anthropic = new Anthropic();

async function answer(question: string, chunks: Chunk[]) {
  const context = chunks
    .map((c, i) => `[${i + 1}] ${c.title} (${c.url})\n${c.content}`)
    .join("\n\n");

  const message = await anthropic.messages.create({
    model: "claude-sonnet-5-5",
    max_tokens: 1024,
    system:
      "Answer using only the numbered sources provided. " +
      "Cite sources like [1]. If the sources don't contain the answer, " +
      "say you don't know and suggest contacting support. Never guess.",
    messages: [
      { role: "user", content: `Sources:\n${context}\n\nQuestion: ${question}` },
    ],
  });

  const block = message.content[0];
  return block.type === "text" ? block.text : "";
}

Two things in that prompt do most of the work. Citations let users check the answer, and they let you check it too. And "say you don't know" turns a confident wrong answer into a handoff to a human, which is what you want.


The parts that actually take the time


The code above is maybe 20% of a real project. The rest is this:


Permissions. Who can see which documents? In most companies the answer is complicated: some docs are public, some are per team, some are per customer. Every chunk needs to carry that information, and every search needs to respect it.


Keeping the index fresh. Documents change. You need a way to notice changes (a webhook, a nightly sync, a "last modified" check), re-chunk and re-embed what changed, and delete chunks for documents that were removed. Stale answers erode trust faster than no answers.


Messy formats. PDFs with tables, scanned documents, slide decks and spreadsheets all need different handling. Tables in particular often come out as scrambled text unless you convert them properly first.


Testing. Before launch, write down 30 to 50 real questions people have actually asked, along with the right answer for each. Run them every time you change the chunking, the prompt or the model. Without this you're tuning by gut feeling.


Learning from misses. Log every question where the bot said "I don't know". That list is the most useful thing the chatbot produces. It tells you exactly which documentation is missing.


Common mistakes


  • Chunks that are too big. The model gets five long passages, only one sentence of which matters, and the answer gets vague.
  • No metadata. Without titles and URLs you can't cite sources, and without permissions you can't launch safely.
  • Letting the model fall back on general knowledge. If it can't find the answer in your docs, it should say so, not fill the gap with something plausible.
  • Skipping evaluation. Everything looks great on the five questions you tried in the demo.

When you don't need RAG at all


If all your content fits comfortably in a single prompt, say a short FAQ or a one-page policy, skip the database and put the whole thing in the system prompt. Modern models handle long context well, and you avoid an entire pipeline.


And if users need the bot to do things, like check an order status or book a meeting, that's not retrieval. That's tool use, where the model calls your API. Often the best assistants combine both: RAG for questions about documents, tools for questions about live data. I wrote about the tool side in MCP explained: how to let AI agents use your app.


Wrapping up


A chatbot that answers from your own documents isn't magic and isn't a research project. It's a search problem with a language model on the end. Get the chunking, permissions and testing right, and the model part is the easy bit.


If you're thinking about adding one to your product or your support site and want help planning it, get in touch. I build AI assistants and knowledge APIs for clients, and I'm happy to tell you honestly whether RAG is the right fit for what you have.

Writing

Notes from building with AI tools, React and production systems.

Get In Touch

Connect

Full Stack Software Engineer passionate about building innovative web and mobile applications.

CEO at Appzivo, the software studio I founded.

© 2026 Saad Minhas. All Rights Reserved.