AI integration services

You already have a product. You want AI inside it, without a rebuild and without a six-month bet.

FROM$4,500
TIMELINE3–7 weeks, one capability
Book a call →
What is AI integration?

AI integration is adding a model-backed capability - search, drafting, classification, an assistant - to software that already exists and already has users. The work is not the model call. It is deciding where AI belongs in the architecture, keeping it inside your existing permission boundaries, proving it works before launch, and making it cheap to turn off.

We do this as a fixed-scope build from $4,500, typically 3 to 7 weeks. You get the service, the evaluation harness, a feature flag, a rollback path and the documentation - and you own all of it. For the full picture of what artificial intelligence integration services involve, read our AI integration services guide.

Who this is for

A product team at a Series A to B company with a working product, paying customers, and an AI feature on the roadmap this quarter. The engineers are strong; none of them has shipped retrieval or an agent to production, and the estimates they give have a very wide error bar because nobody knows yet where the hard parts are.

If you would rather have that expertise permanently on the team, a dedicated engineer is the cheaper shape. If there is no product yet, start with AI product development.

Where AI belongs in your architecture

  1. 01A separate service, not a library inside your monolith. Model calls are slow and occasionally fail in ways your existing error handling was not written for. Isolating them keeps a provider outage from becoming an outage of your product.
  2. 02Behind your existing auth. The AI service never talks to a browser directly and never holds its own user model. It receives an already-authenticated identity and applies it.
  3. 03Its own storage for embeddings and traces. Kept separate so you can re-index, re-embed or delete without touching production data.
  4. 04A queue for anything that is not interactive. Bulk classification and document ingestion do not belong in a request-response path.
  5. 05A flag in front of the whole thing. Off by default, on per account, so a rollout is a configuration change.
Model choice

Choosing a model, with the arithmetic

We benchmark two or three candidates against your own evaluation set in the first week, and the decision comes from that table rather than from a vendor's benchmark. The pattern that wins most often is a small model on the hot path with a larger one behind a confidence threshold.

How model routing affects relative cost and p95 latency
Routing strategyRelative costp95 latencyWhen it wins
Large model on every request1.0x3-6sLow volume, high stakes, answers a human reviews anyway
Small model, large as fallback0.25-0.4x1-2s typical, 4s on fallbackMost production features. The default we reach for.
Small model plus a re-ranker0.3-0.5x1-2sRetrieval quality is the bottleneck, not generation
Cached answers for repeated queries0.05-0.2x on cache hitsunder 200msSupport and documentation search, where queries repeat heavily

Costs are relative to the same feature built on a frontier model for every request; absolute numbers depend on your token volume and cache hit rate. We measure both in week one.

Auth and data boundaries

01

Permissions live in retrieval, not in the prompt

Asking a model not to reveal something is not a security control. The user's identity is carried into the retrieval query, so documents they cannot see are never candidates and never reach the context window. Every retrieval is filtered before ranking, not after.

02

Tenant isolation, where it applies

Separate namespaces per tenant in the vector store, separate encryption scope, and a test in the suite that asserts a query from tenant A cannot return a document from tenant B. That test runs on every deploy. We built exactly this for the Klebbix platform.

03

What leaves your infrastructure

Only what the model needs, and only to providers you have approved in writing. No client data is used to train a model, ours or a vendor's - the relevant provider settings are configured and documented in handover. The full subprocessor list is on the security page.

Process

Evaluation before launch

  1. 01week 1

    Collect real queries

    Fifty to two hundred real inputs from your support tickets, search logs or user interviews. Not invented examples - invented examples are always easier than reality.

  2. 02week 1

    Label what good looks like

    Someone who knows the domain marks the correct answer, or the acceptable range of answers. This is the expensive hour and it is not optional.

  3. 03week 1

    Set thresholds before building

    Agree what score ships and what does not, in writing, before anyone is emotionally invested in the implementation.

  4. 04ongoing

    Run on every change

    Prompt edits, model swaps, chunking changes and retrieval tweaks all run the suite. Non-determinism is handled by running each case several times and reporting the spread, not the best result.

  5. 05yours

    Keep it after handover

    The harness ships with the build and the documentation explains how to extend it. It is the part of the delivery with the longest useful life.

Rollback strategy

Every integration we ship has three ways back, because the failure you plan for is never the one you get.

01

Feature flag

A feature flag that disables the AI path per account without a deploy.

02

Pre-AI behaviour kept

The pre-AI behaviour kept intact and reachable, not deleted on launch day.

03

Logged traces

A logged trace of retrieval, prompt and response for every request, so a bad output can be reproduced rather than argued about.

The stack

The stack we usually reach for

Named, because vagueness here is what separates content-mill agencies from people who have shipped. We will use yours instead wherever you already have a preference.

Default stack by layer for an AI integration
LayerDefault
ModelsOpenAI and Anthropic APIs, with Cohere Rerank for retrieval quality.
Vector storageQdrant self-hosted, or pgvector when the corpus is small enough that a second datastore is not worth operating.
OrchestrationPlain application code by default; LangGraph or Temporal when a workflow genuinely needs durable state.
CachingRedis, keyed on a normalised query plus the permission scope.
EvaluationA project-local harness with a labelled set, run in CI.
ObservabilityTraces per request, cost per request, and Grafana dashboards for both.

Deeper detail on the retrieval half of this is on the RAG systems page, and worked examples are in the case studies.

What it costs

From $4,500 for one well-defined capability. Most integrations land between $6,000 and $18,000 depending on how many systems the feature has to touch and whether the data needs work first.

Fixed price, fixed date, quoted after a scoping call - see pricing for what moves the number.

Frequently asked questions

Send us the feature and we will scope it

A free strategy session and a fixed quote.