White-label RAG and document search
The project agencies resell most, and the one with the widest margin.
We build retrieval-augmented document search under your agency's brand: hybrid search across your client's contracts, manuals, tickets or knowledge base, with answers cited back to the source page and permissions enforced inside the retrieval query.
Our cost is $5,500 to $9,000 for one corpus, against a typical resale of $14,000 to $22,000 - a margin near 58%. Five to seven weeks, fixed scope and fixed price, delivered in your repositories with an NDA and non-solicit signed before scoping.
The economics
| Scope | You quote | Our cost | Your margin | Timeline |
|---|---|---|---|---|
| Search over one corpus, with citations | $14,000-22,000 | $5,500-9,000 | ~58% | 5-7 weeks |
| Multi-tenant, with isolation and SSO | $26,000-40,000 | $12,000-20,000 | ~52% | 8-12 weeks |
| Adding retrieval to an existing product | $10,000-16,000 | $4,500-7,000 | ~56% | 4-6 weeks |
| Corpus audit and feasibility report | $4,000-6,000 | $2,000-3,000 | ~52% | 1-2 weeks |
The last row is worth selling on its own. A paid two-week audit de-risks the build for your client, gets you paid for discovery, and means the fixed quote that follows is one we can both stand behind.
What is included in every build
- 01An ingestion pipeline for the client's real formats - including the scanned PDFs from before 2019 that nobody mentions at scoping.
- 02Hybrid retrieval: vector search for meaning, structured filters for dates, codes and identifiers, and a re-ranker for precision.
- 03Answers cited to the source document and page, with an explicit refusal when retrieval finds nothing.
- 04Permission filtering inside the retrieval query, and per-tenant namespaces with an automated isolation test where it applies.
- 05An evaluation harness: a labelled query set from the client's real questions, regression runs, agreed thresholds.
- 06Cost and latency instrumentation, plus an admin interface someone non-technical can actually use.
The engineering behind each of these is described on our RAG systems page.
How to scope it with your client
Four questions that turn a vague brief into something quotable. Ask them before the scoping call and we can usually price in one pass.
- 01Which documents, and in what formats?Ask for ten real files, including the worst ones. A clean sample and the real archive are different projects.
- 02What questions do people ask today?Twenty real ones. These become the evaluation set, and they are also the honest test of whether retrieval is the right answer at all.
- 03Who is allowed to see what?If the answer is "everyone sees everything", the build is materially simpler. If not, permissions shape the architecture from day one.
- 04How often do documents change?A static archive and a corpus that updates hourly need different ingestion, and the difference is real money.
What not to promise
- 01"It will answer anything."It answers what is in the documents. If the knowledge lives in people's heads, retrieval will expose that rather than fix it - which is useful, but it is not what was sold.
- 02An accuracy number before seeing the corpus.Any figure quoted at that stage is invented. Sell the evaluation step instead: we measure it in week one and you report a real number.
- 03"It replaces the intranet search."Often true eventually, but the first release should be scoped to one corpus and one team. Wide first releases are how these projects lose their deadline.
Proof
The Klebbix hybrid retrieval system - unified retrieval across spreadsheets, documents and a relational database for a European SaaS provider, multi-tenant, GDPR and ISO-27001 scoped.
Published under our name, so it is not yours to forward. Ask us for an anonymised technical summary to attach to a proposal and we will write one.
How the engagement runs
- 01The processNDA first, scoping call within two working days, fixed quote within three, delivery in your repositories under an identity you control. The detail is on how a white-label engagement works, and the invisibility commitment is on the white-label page.
- 02Need an agent instead?If the client wants a system that acts rather than answers, see white-label AI agents.
Frequently asked questions
$14,000 to $22,000 for retrieval over one corpus with citations and an admin interface, against our cost of $5,500 to $9,000. Multi-tenant with isolation and SSO quotes at $26,000 to $40,000 against our $12,000 to $20,000. Both leave a margin around 58%.
Ten representative documents, including the worst ones, and the questions people actually ask. Nothing changes an estimate more than seeing the real corpus - a clean sample and the real archive are different projects, and we would rather find that out before you have quoted.
On a well-defined corpus with a labelled test set, 90% or better on retrieval relevance is a reasonable target; we measured 93%+ on the Klebbix build. We will not give you a number to put in a proposal before seeing the documents, because any number given at that stage is invented.
Yes, and it is built in rather than bolted on. Permission filtering happens inside the retrieval query, so a user cannot retrieve what they cannot see, and multi-tenant builds get separate namespaces plus an automated test asserting one tenant's query cannot return another's document.
$150 to $600 a month at moderate volume for embeddings, vector storage and inference, on their own accounts at cost. Quote it separately and honestly - a running cost that appears as a surprise in month two is the fastest way to sour an otherwise successful project.
Yes, and you should. Corpora grow, documents change format, and the eval set needs extending as new question types appear. A quarterly review or a retained few hours a month is a straightforward margin on top of the build.
Send ten of your client's documents
That plus twenty real questions is enough for a fixed quote within three working days.
