RAG development services
Answers grounded in your own documents, with citations, at a latency people will tolerate.
A retrieval-augmented generation system finds the right passages in your own documents and databases, then has a model answer from them with a citation back to the source. That is what makes an answer checkable instead of merely fluent.
We build these as hybrid systems - vector search for meaning, structured filters for identifiers and dates, a re-ranker for precision - with per-tenant isolation where it is needed and an evaluation harness that proves the thing still works after a change. From $5,500, 5 to 7 weeks for one corpus.
Who this is for
Teams whose people spend real hours searching: support agents hunting through documentation, operations staff reading contracts, analysts reconciling reports that live in four systems. The test for whether RAG is worth building is simple - if the answer exists in your documents and someone is paid to go and find it, retrieval pays for itself quickly.
What you get
- 01An ingestion pipeline that handles your real formats - including the scanned PDFs from before 2019 that nobody mentioned at scoping.
- 02Hybrid retrieval: vector search over passages, structured filters over metadata, and a re-ranker on top.
- 03Answers with citations back to the source document and page, and a refusal path when retrieval finds nothing.
- 04Permission filtering inside the retrieval query, so a user cannot retrieve what they cannot see.
- 05An evaluation harness: a labelled query set, regression runs on every change, agreed pass thresholds.
- 06Cost and latency instrumentation per request, with dashboards.
- 07Documentation, runbooks and a recorded walkthrough at handover.
You do not need an engineering team to start
Describe the job you want done in plain words - answering customer questions, sorting incoming requests, pulling figures out of documents. On the first call we tell you whether AI is the right tool for it, roughly what it would cost and how long it would take.
If it is not worth building, we say so. If you are building a whole product rather than one tool, read how an AI product build works.
- Single-purpose agent$4,500-8,000
An assistant that answers your customers' questions from your own help pages and product documents, with a link to where each answer came from.
Typically 2-4 weeks - Production agent$8,000-18,000
An assistant that reads each incoming request, looks the customer up in your systems, drafts the reply and waits for a person on your team to approve it before anything is sent.
Typically 4-7 weeks
How we build it
- 01week 1
Corpus audit and eval set
We look at the real documents, not a sample someone tidied. Formats, volume, update frequency, and the parts that are unreadable. In parallel, 50-200 real questions get labelled with correct answers - this defines 'working' before anyone writes retrieval code.
- 02weeks 2-3
Ingestion and chunking
Parsing per format, OCR where needed, and a chunking strategy chosen by measurement rather than convention. Chunk size and overlap are tuned against the eval set; the default that works for everyone does not exist.
- 03week 4
Hybrid retrieval
Embeddings in Qdrant or pgvector, structured fields in Postgres, and an orchestrator that merges both into one ranked context. Cohere Rerank on the candidate set, because precision at the top of the list is what the model actually reads.
- 04week 5
Generation, citation and refusal
Answer synthesis with source citations, confidence thresholds, and an explicit 'I could not find this' path. A system that always answers is a system that sometimes invents.
- 05week 6
Isolation, caching and hardening
Per-tenant namespaces and the isolation test, Redis caching on normalised queries within a permission scope, rate limits and audit logging.
- 06week 7
Evaluation and handover
Full run against the week-1 set, load test, cost-per-query measurement, documentation and a recorded walkthrough.
The default RAG stack, layer by layer
We will use yours instead wherever you already have a preference.
| Layer | Default | Why |
|---|---|---|
| Vector store | Qdrant self-hosted, or pgvector | Qdrant for per-tenant namespaces and payload filtering at scale; pgvector when the corpus is small enough that a second datastore is not worth operating. |
| Structured data | PostgreSQL | Metadata, permissions and the filters that embeddings are bad at - dates, carriers, identifiers. |
| Re-ranker | Cohere Rerank | Cheap precision. Retrieve broadly, re-rank narrowly, send less to the model. |
| Embeddings | OpenAI or Cohere, chosen by measurement | Benchmarked on your corpus in week one rather than picked from a leaderboard. |
| Cache | Redis, keyed on normalised query + permission scope | The largest single lever on running cost. |
| Serving | FastAPI or Node, in Docker | Boring, observable, and deployable into your existing infrastructure. |
The Klebbix hybrid retrieval system is this service at production scale: unified retrieval across spreadsheets, documents and a relational database for a European SaaS provider, with full tenant isolation under GDPR and ISO-27001 scope.
From $5,500 for retrieval over one corpus, $12,000 to $20,000 for multi-tenant with isolation and SSO. Fixed price and fixed date after a scoping call. Running cost is typically $150 to $600 a month at moderate volume, on your own accounts at cost.
See pricing for the full rate card. If you are an agency reselling this, white-label RAG has the margin maths.
Frequently asked questions
Retrieval-augmented generation: instead of hoping a model already knows something, you retrieve the relevant passages from your own documents and give them to the model along with the question. The model's job becomes reading and summarising rather than remembering, which is what makes citations and up-to-date answers possible.
Five to seven weeks for retrieval over one corpus, eight to twelve if it is multi-tenant with isolation requirements and SSO. The first two weeks are almost entirely about your documents, not about models.
On a well-defined corpus with a labelled test set, 90% or better on retrieval relevance is a reasonable target - we measured 93%+ on the Klebbix build. Anyone quoting an accuracy number without naming the test set is quoting a number that means nothing.
Both, almost always. Vector search finds things phrased differently from the query; keyword and structured filters find exact identifiers, dates and codes that embeddings routinely miss. Hybrid retrieval plus a re-ranker beats either alone on every corpus we have measured.
pgvector if your corpus fits comfortably and you already run Postgres - one less system to operate is worth a great deal. Qdrant self-hosted when you need per-tenant namespaces, payload filtering at scale or millions of vectors. We do not have a favourite to sell you; the cost and operational burden of each is on our vector database comparison.
Separate namespaces per tenant, permission filters applied inside the retrieval query rather than after it, and an automated test that asserts a query from tenant A cannot return a document from tenant B. That test runs on every deploy, because isolation you do not test is isolation you are hoping for.
$150 to $600 a month for a typical single-corpus system at moderate volume, covering embeddings, vector storage and inference, billed to your own accounts at cost. Caching repeated queries is the single largest lever - a good cache hit rate cuts inference spend by more than any model change.
Send us ten real documents
Ten representative files and the questions people ask about them. That is enough for a fixed quote.
