RAG
Private AI over your own data
Answers from your own records, with the source attached.
The short version
Search and answering across your internal documents, contracts, records and history — running where your data is allowed to be.
Your organisation already holds the answers to most of the questions your staff ask each other. They are in ten years of documents, in a case management system, in a shared drive nobody has indexed. Retrieval-based systems make that searchable in plain language without your data leaving your control.
What is included
Concretely, this is the work.
- Document ingestion and chunking across PDFs, scans, spreadsheets, email archives and case systems
- Vector search with hybrid keyword matching, because pure semantic search misses names, codes and reference numbers
- Answers that cite the source passage every time, so a person can verify rather than trust
- Permission-aware retrieval — a user can only get an answer from documents they were already allowed to open
- Self-hosted deployment with open-weight models where data cannot leave your infrastructure
- Redaction and PII handling before anything reaches a model
- Evaluation against a real question set, so you know the accuracy rate rather than guessing at it
Approach
What you should expect to change
- Questions answered in seconds that previously meant asking a colleague or reading for an hour
- Institutional knowledge that survives the person who held it leaving
- A measured accuracy figure, with citations, rather than a vague sense that it seems to work
How we build private AI systems
The hard part is not the model. It is retrieval — finding the six paragraphs out of four million that actually answer the question. Most disappointing internal AI projects fail here, and no amount of model quality rescues bad retrieval.
So we spend the first phase on your corpus: how it is structured, how it is named, what the reference numbers mean, which documents supersede which. Then we build hybrid retrieval and measure it against questions your staff genuinely asked last month.
On data residency
Where a commercial API is contractually acceptable, it is usually the faster and cheaper answer. Where it is not — and in health, legal and government work it frequently is not — we deploy open-weight models on infrastructure you control, and the data never crosses the boundary. We will tell you plainly which situation you are in rather than selling you the more expensive one.
Questions about private AI over your own data
Does our data get used to train anything?
Not in anything we build. With self-hosted models the question does not arise. With commercial APIs we use enterprise terms that exclude training, and we tell you exactly which provider handles what.
How accurate is it?
That depends on your corpus and we will not pretend otherwise. What we will do is measure it — build a set of real questions with known answers, score the system against it, and show you the number before you commit to a rollout.
Also
Other things we build.
AI agents and workflow automation
Agents that carry a whole process end to end — reading the enquiry, checking availability, drafting the quote, escalating the parts a person should see.
Read more BKGBooking and reservation platforms
The system your revenue actually passes through — search, availability, quoting, payment, confirmation, amendment, cancellation. Built to survive the edge cases.
Read more OPSOperations and back-office systems
The platform your team lives in all day. Suppliers, jobs, staff, documents, margins, reporting — shaped around how you actually work rather than a generic template.
Read moreStart a conversation
Have a project like this?
Send us a paragraph describing what the system needs to do and what you are using now. We will come back with an honest view of whether it is a fit and roughly what it costs.