AI integration services
You already have a product. You want AI inside it, without a rebuild and without a six-month bet.
AI integration is adding a model-backed capability - search, drafting, classification, an assistant - to software that already exists and already has users. The work is not the model call. It is deciding where AI belongs in the architecture, keeping it inside your existing permission boundaries, proving it works before launch, and making it cheap to turn off.
We do this as a fixed-scope build from $4,500, typically 3 to 7 weeks. You get the service, the evaluation harness, a feature flag, a rollback path and the documentation - and you own all of it. For the full picture of what artificial intelligence integration services involve, read our AI integration services guide.
Who this is for
A product team at a Series A to B company with a working product, paying customers, and an AI feature on the roadmap this quarter. The engineers are strong; none of them has shipped retrieval or an agent to production, and the estimates they give have a very wide error bar because nobody knows yet where the hard parts are.
Where AI belongs in your architecture
- 01A separate service, not a library inside your monolith. Model calls are slow and occasionally fail in ways your existing error handling was not written for. Isolating them keeps a provider outage from becoming an outage of your product.
- 02Behind your existing auth. The AI service never talks to a browser directly and never holds its own user model. It receives an already-authenticated identity and applies it.
- 03Its own storage for embeddings and traces. Kept separate so you can re-index, re-embed or delete without touching production data.
- 04A queue for anything that is not interactive. Bulk classification and document ingestion do not belong in a request-response path.
- 05A flag in front of the whole thing. Off by default, on per account, so a rollout is a configuration change.
Choosing a model, with the arithmetic
We benchmark two or three candidates against your own evaluation set in the first week, and the decision comes from that table rather than from a vendor's benchmark. The pattern that wins most often is a small model on the hot path with a larger one behind a confidence threshold.
| Routing strategy | Relative cost | p95 latency | When it wins |
|---|---|---|---|
| Large model on every request | 1.0x | 3-6s | Low volume, high stakes, answers a human reviews anyway |
| Small model, large as fallback | 0.25-0.4x | 1-2s typical, 4s on fallback | Most production features. The default we reach for. |
| Small model plus a re-ranker | 0.3-0.5x | 1-2s | Retrieval quality is the bottleneck, not generation |
| Cached answers for repeated queries | 0.05-0.2x on cache hits | under 200ms | Support and documentation search, where queries repeat heavily |
Costs are relative to the same feature built on a frontier model for every request; absolute numbers depend on your token volume and cache hit rate. We measure both in week one.
Auth and data boundaries
Permissions live in retrieval, not in the prompt
Asking a model not to reveal something is not a security control. The user's identity is carried into the retrieval query, so documents they cannot see are never candidates and never reach the context window. Every retrieval is filtered before ranking, not after.
Tenant isolation, where it applies
Separate namespaces per tenant in the vector store, separate encryption scope, and a test in the suite that asserts a query from tenant A cannot return a document from tenant B. That test runs on every deploy. We built exactly this for the Klebbix platform.
What leaves your infrastructure
Only what the model needs, and only to providers you have approved in writing. No client data is used to train a model, ours or a vendor's - the relevant provider settings are configured and documented in handover. The full subprocessor list is on the security page.
Evaluation before launch
- 01week 1
Collect real queries
Fifty to two hundred real inputs from your support tickets, search logs or user interviews. Not invented examples - invented examples are always easier than reality.
- 02week 1
Label what good looks like
Someone who knows the domain marks the correct answer, or the acceptable range of answers. This is the expensive hour and it is not optional.
- 03week 1
Set thresholds before building
Agree what score ships and what does not, in writing, before anyone is emotionally invested in the implementation.
- 04ongoing
Run on every change
Prompt edits, model swaps, chunking changes and retrieval tweaks all run the suite. Non-determinism is handled by running each case several times and reporting the spread, not the best result.
- 05yours
Keep it after handover
The harness ships with the build and the documentation explains how to extend it. It is the part of the delivery with the longest useful life.
Rollback strategy
Every integration we ship has three ways back, because the failure you plan for is never the one you get.
Feature flag
A feature flag that disables the AI path per account without a deploy.
Pre-AI behaviour kept
The pre-AI behaviour kept intact and reachable, not deleted on launch day.
Logged traces
A logged trace of retrieval, prompt and response for every request, so a bad output can be reproduced rather than argued about.
The stack we usually reach for
Named, because vagueness here is what separates content-mill agencies from people who have shipped. We will use yours instead wherever you already have a preference.
| Layer | Default |
|---|---|
| Models | OpenAI and Anthropic APIs, with Cohere Rerank for retrieval quality. |
| Vector storage | Qdrant self-hosted, or pgvector when the corpus is small enough that a second datastore is not worth operating. |
| Orchestration | Plain application code by default; LangGraph or Temporal when a workflow genuinely needs durable state. |
| Caching | Redis, keyed on a normalised query plus the permission scope. |
| Evaluation | A project-local harness with a labelled set, run in CI. |
| Observability | Traces per request, cost per request, and Grafana dashboards for both. |
Deeper detail on the retrieval half of this is on the RAG systems page, and worked examples are in the case studies.
From $4,500 for one well-defined capability. Most integrations land between $6,000 and $18,000 depending on how many systems the feature has to touch and whether the data needs work first.
Fixed price, fixed date, quoted after a scoping call - see pricing for what moves the number.
Frequently asked questions
Almost never. In most codebases the AI work lands as a new service behind your existing API, plus a feature flag and an evaluation harness. We touch your product code as little as possible, because the parts of it that already work are not the risk.
Usually the cheapest one that passes your evaluation set, with a larger one behind a fallback for the hard cases. Model choice is rarely the expensive decision - integration count and data quality are. We benchmark two or three candidates on your actual data in week one and let the numbers pick.
Retrieval is filtered by the same permissions your application already enforces, before anything reaches the model, not after. That means the user's identity is carried into the retrieval query rather than the prompt, so an unauthorised document cannot be retrieved in the first place.
A labelled evaluation set built from your real queries, run on every change, with agreed pass thresholds. We build it in week one, before any feature work, because a system with no eval set cannot be improved - only changed.
Three things, and we build all of them: a feature flag that turns the AI path off without a deploy, a logged trace of the retrieval and prompt for every response, and a fallback to the non-AI behaviour your product had before. Rollback is a configuration change, not a release.
Cache aggressively, route easy requests to a smaller model, cap tokens per request, and alert on cost per user rather than total spend. We instrument this from the first week - a system that is cheap in testing and expensive in production is usually a system nobody measured per-request.
Three to seven weeks for one well-defined capability in an existing product, longer if the data needs work or several systems are involved. We give a fixed price and a fixed date after a scoping call.
Send us the feature and we will scope it
A free strategy session and a fixed quote.
