Who it's for

  • Teams whose knowledge is scattered across docs, PDFs and wikis
  • Support teams answering the same questions every day
  • Companies that want an AI assistant without sending everything to a black box

What you get

  • Ingestion of your documents with a plan for keeping them up to date
  • Retrieval tuned on your real questions, with cited sources in every answer
  • A chat interface for your team or your website
  • An evaluation set so answer quality can be measured, not guessed
  • Clear notes on costs, data handling and limits

How a RAG chatbot works

Your documents are split into passages, indexed for semantic search, and searched for each question. The LLM answers only from what it finds and links back to the source, so people can check it. See the full pipeline in the diagram below.

RAG or fine-tuning?

For most company chatbots, RAG is the right choice: your documents change, people need sources, and nothing has to be retrained when you add a file. Fine-tuning is for teaching a model a format or tone, not for storing facts. More in RAG or fine-tuning?

Built to be measured

Before launch we collect real questions from your team or customers and use them as a test set. Changes to prompts, models or documents are checked against it, so quality goes up rather than drifting.

Proof

Real products I built for real clients: the problem, what I built, and the stack.

PayRight validation report for an invoice, with accuracy score, invoiced and expected totals, variance, potential savings and linked contract and purchase order

RAG and AI chatbots

PayRight: AI that checks every supplier invoice against the contract, with a human in the loop

A UK contract compliance and invoice assurance platform. AI extracts contract clauses, people confirm them, and every invoice is scored against contract and PO.

Client
PayRight, powered by GovernTerms (UK)
Role
AI full-stack developer
When
2026
  • Claude Code
  • React
  • Node.js + TypeScript
  • PostgreSQL / Supabase
  • OCR and document extraction
  • Role-based access control
StoreFilter chat answering a question about a Shopify store's conversion rate, with the calculation, insights and follow-up questions

RAG and AI chatbots

StoreFilter: an AI analyst for any Shopify store, built on 2.5 million records

An AI analytics chat for e-commerce, built with Next.js, Claude Code, Supabase, pgvector and GPT-4. Ask about any Shopify store, get answers with charts.

Client
StoreFilter
Role
Full-stack developer
When
2025
  • Next.js
  • Claude Code
  • Supabase
  • PostgreSQL + pgvector
  • Edge Functions
  • OpenAI GPT-4

How it's built

The architecture, step by step

An assistant that answers from your own documents

  1. Your documents

    PDFs, docs, wikis, help centre

  2. Chunk and embed

    Passages turned into vectors

    • Python
  3. Vector index

    Semantic search over passages

  4. Retrieve

    The best passages per question

  5. LLM

    Answers only from what was found

    • Claude
    • GPT
  6. Cited answer

    Links back to the source

Why this stack

  • No model training on your data: the assistant reads your documents at question time.
  • The LLM is chosen per project on answer quality, cost and your data rules.
  • Every answer shows its sources, so people can check it.

An evaluation set of real questions is re-run on every change.

Typical starting points. The final stack is chosen per project and written into your scope.

Related reading

Questions

Will it make things up?

Every answer is grounded in retrieved passages and shows its sources, and we test against a set of your real questions. When the documents don't contain the answer, it says so.

Do you need to train a model on our data?

Usually not. A RAG chatbot reads your documents at question time, so there is no model training and documents can be added or removed at any time.

Can the chatbot go on our website?

Yes. The same assistant can serve your team internally or your customers on your website, with different document sets for each.

Which LLM do you use?

It's chosen per project based on answer quality, cost and your data rules, typically Claude (Anthropic) or GPT (OpenAI).

Show me the documents your team keeps searching

A 30-minute call is enough to tell whether it fits, what the first version should include, and how long it takes.

Book a 30-minute call