How-toby Ahsan Ahmad

AI integration services: adding AI to software you already run

Summary

AI integration services add capabilities such as search over your documents, drafting, classification and extraction to software you already run, without rebuilding it. The work usually lands as a new service behind your existing API, with retrieval filtered by your existing permissions, a feature flag, a fallback to current behaviour and an evaluation set built from your real data. Calling a GPT or Claude API is the easy part; the integration around it is the job. We price integration from $4,500 for one capability, with most landing between $6,000 and $18,000 over three to seven weeks.

What are AI integration services?

AI integration services add AI capabilities - search over your documents, drafting, classification, extraction, recommendations - to software you already run, without rebuilding it. In most codebases the work lands as a new service behind your existing API, wired into your data and permissions, with a feature flag in front of it and a test suite that tells you whether it still works after the next change.

That is different from building an AI product from nothing, which is AI product development. Integration is usually the cheaper and faster job, because the product, the users and the data already exist. The hard part is fitting a probabilistic component into a system that was built to behave deterministically.

Is GPT development the same as AI integration?

Mostly, yes. "GPT development" usually means building features on top of a large language model API - OpenAI's GPT models, Anthropic's Claude or an open-weight model - rather than training a model. Calling the API is the easy part and takes an afternoon. The integration work is everything around the call:

  • Getting the right context into the prompt from your own data, filtered by what the user is allowed to see
  • Structuring the output so the rest of your code can rely on it
  • Handling timeouts, rate limits, refusals and wrong answers without breaking the page
  • Measuring quality with a labelled test set, and cost per request, from the first week
  • A fallback to the product's old behaviour when the model is unavailable or unsure

Model choice matters less than people expect. We usually benchmark two or three candidates on your actual data and pick the cheapest one that passes the evaluation set, with a larger model behind a fallback for the hard cases.

What does artificial intelligence systems integration involve?

Artificial intelligence systems integration is the engineering that connects a model to the systems around it: your database, your document stores, your identity provider, your CRM or ERP, and your deployment pipeline. It is where most integration projects succeed or fail.

Data access

A model is only as useful as the context it gets. For questions over your own content that means a retrieval system: documents chunked, indexed and searched, with results passed to the model. For structured data it means tools the model can call, such as a function that looks up an order.

Permissions

Retrieval has to be filtered by the same permissions your application already enforces, before anything reaches the model. The user's identity belongs in the retrieval query, not in the prompt, so a document they cannot see is never retrieved in the first place.

Operations

A logged trace of the retrieval and prompt for every response, cost and latency per request, alerts, and a rollback that is a configuration change rather than a release. Without these, the first bad answer in production turns into a guessing game.

Which AI features are most commonly integrated into existing software?

Five patterns account for most of the integration work we see. They differ mainly in how much data work they need and how bad a wrong answer is.

Common AI integration patterns, what they need and where they go wrong
PatternWhat it needsMain risk
Search and answers over your contentRetrieval index, permission filtering, citationsConfident answers from the wrong document
Drafting and summarisingGood templates, context from the record, human reviewTone or facts that a user sends without checking
Classification and routingA labelled history to test againstSilent drift as categories change
Extraction from documentsValidation against known records, a review queuePlausible but wrong field values
In-product assistants and agentsTools with narrow permissions, approval stepsActions taken that should have been asked about

If the feature takes actions rather than just answering, it is closer to agent development, and the approval and permission design deserves more of the budget.

Do you need a machine learning app development company or an LLM integrator?

It depends on the problem. If you are predicting a number or a category from structured, historical data - churn, demand, fraud scores - you want classical machine learning, and a team that trains and evaluates models on your data. If the problem involves language, documents or open-ended requests, a pre-trained language model integrated into your product usually beats training anything yourself.

Most machine learning app development today is the second kind. Fine-tuning or training your own model makes sense in narrower cases: very high volume where a smaller tuned model is much cheaper to run, strict latency or on-premise requirements, or a task where the general models consistently fail your evaluation set. Try retrieval and prompt design first; they are cheaper to change.

How does an AI integration project run?

Our AI integration projects take three to seven weeks for one well-defined capability in an existing product, longer if the data needs work or several systems are involved. The order matters more than the length:

  • Scope one capability.One feature, one user, one success measure written as a number. "Add AI" is not a scope.
  • Build the evaluation set first. Labelled examples from your real queries or documents, and pass thresholds agreed before implementation. We explain why in the eval harness we ship with every build.
  • Benchmark models on your data. Two or three candidates, scored against the set, with cost per request alongside quality.
  • Integrate behind a flag. A new service behind your existing API, permission-filtered, with a fallback to current behaviour.
  • Roll out gradually. Internal users, then a slice of customers, with traces and cost dashboards watched throughout.

What goes wrong in AI integration projects?

Most failed integrations fail for reasons that have little to do with the model. These are the patterns we see most often when we are brought in to rescue one:

No way to tell whether it works

The team tried the feature on a dozen questions, it looked good, and it shipped. Then a prompt change fixed one complaint and quietly broke three other cases nobody was checking. Without a labelled evaluation set, every change is a gamble, and the team ends up afraid to touch the prompt at all.

Permissions added afterwards

The prototype had access to every document because that was easiest. Adding permission filtering later means rebuilding the retrieval layer, and until it is rebuilt the feature cannot go to customers. Designing permissions in from the first week is far cheaper.

Data that was assumed to be clean

The knowledge base turns out to hold three contradictory versions of the same policy, or half the documents are scanned images with no text layer. The model then answers confidently from the wrong source. Data work is often the largest single item in an integration budget, and it should be estimated up front rather than discovered.

Cost nobody measured

A feature that costs fractions of a cent in testing can cost real money at production volume, particularly when every request stuffs a long context into a large model. Measuring cost per request from the start makes this visible while it is still cheap to fix.

How do you keep an AI integration working after launch?

An integrated AI feature needs more upkeep than ordinary code, because its inputs keep changing. New documents arrive, users ask questions nobody anticipated, and model providers update or retire the models you depend on. Three habits keep it healthy:

  • Run the evaluation set on every change to prompts, retrieval or model version, and block releases that fall below the agreed thresholds
  • Sample real production traces every week, label the failures, and add them to the evaluation set so the same mistake is caught next time
  • Watch cost and latency per request as closely as error rates, and set alerts on both

The handover should leave your own team able to do all three. If ongoing work continues, a dedicated engineer embedded in your team is usually the most economical way to keep improving the feature.

How much do artificial intelligence integration services cost?

We price integration from $4,500 for one well-defined capability. Most integrations land between $6,000 and $18,000, depending on how many systems the feature has to touch and whether the data needs work first. Model API spend and infrastructure go on your own accounts at cost.

The cost drivers, in order of how much they usually move the number: data quality and access, the number of systems involved, how much a wrong answer costs (which decides review and testing effort), and only then the model. A fuller breakdown is in what AI development costs, and our rates are on the pricing page.

Keeping running cost predictable is part of the build: cache aggressively, route easy requests to a smaller model, cap tokens per request, and alert on cost per user rather than total spend.

How do you choose an AI integration services provider?

Look for a team that talks about your data, your permissions and how they will measure quality before they talk about models. Specific things to check:

  • They propose an evaluation set built from your data, with thresholds agreed in advance
  • They explain how retrieval respects your existing permissions
  • They plan a feature flag, a fallback and per-request tracing from the start
  • They touch your existing product code as little as possible
  • You own the code, prompts and evaluation data from the first commit
  • They can show a production system they built, with real numbers

Two of ours are written up in full on the case studies page. If you have a capability in mind, book a call and we will tell you what it would take and whether it is worth doing.

Tell us what you are building

A free strategy session: what is buildable, what it costs, what we would not attempt.