When to use LangChain
A framework is a bet that your problem resembles the one it was built for. Sometimes that bet is right.
By Ahsan Ahmad, Chief Executive Officer · Last reviewed
Use LangGraph when a workflow needs durable state - branching, checkpoints, human-in-the-loop pauses, or resumption after a process restart. That is a real problem and building it yourself is a month you did not plan for.
Write it yourself when the system is a retrieval pipeline or a short agent loop. Embed, search, re-rank, prompt is four steps and a couple of hundred lines you fully understand - and understanding exactly what was sent to the model is where production debugging time actually goes.
When is LangChain worth using?
When a workflow needs durable, resumable state, when you need many of its document loaders and connectors, when you are proving an approach in an afternoon, or when your team already knows and maintains it. Outside those cases, a retrieval pipeline or a short agent loop is usually clearer written directly against the provider SDK.
- Durable, resumable workflows. LangGraph's checkpointing means a workflow can pause for a human approval on Tuesday and continue on Wednesday. Writing that yourself is a real project.
- Breadth of integrations. Dozens of document loaders and connectors that are tedious rather than difficult to write. If you need eleven of them, the ecosystem is the reason to be there.
- Prototyping speed. Getting to a working demonstration in an afternoon has genuine value when the question is whether an approach works at all.
- A team that already knows it. Familiarity usually beats architectural preference. We would rather inherit a LangChain codebase a team can maintain than a bespoke one they cannot.
What does LangChain cost you in production?
Mostly debugging distance, not speed. Framework overhead is milliseconds against model calls measured in seconds, but every layer between your code and the model's HTTP call is a layer to see through when something breaks. You also take on upgrade risk from breaking changes you did not choose, and a framework to learn before the code.
| Framework | Direct implementation | |
|---|---|---|
| Time to first demo | Hours | A day or two |
| Seeing the exact prompt sent | Through a tracing layer | It is in the code |
| Upgrade risk | Breaking changes you did not choose | Provider SDK only |
| Onboarding a new engineer | Learn the framework, then the code | Learn the code |
| Unusual requirements | Fight or fork the abstraction | Write it |
| Durable multi-step state | Solved for you | A month of work |
| Debugging a production failure | Peel back layers first | Read the function |
The row that matters most is the last one. Every production investigation of an LLM system ends at the same question - what exactly went to the model, and what came back - and each layer between your code and that HTTP call is a layer to see through first. Good tracing makes this manageable; the absence of it is where the hours go.
What we use, and why
- Provider SDKs directly for the model calls. One dependency, documented by the people who run the API, and the request is visible in the code.
- Plain functions for retrieval. Embed, search, filter, re-rank, assemble context. Each step testable on its own, which is what makes the evaluation harness possible.
- LangGraph when durable state is a genuine requirement - long-running workflows with approval steps.
- Temporal when durability matters more than the LLM part, which happens more often than people expect once a workflow touches payments or external commitments.
- A tracing layer, always. Whatever the framework choice, every request stores the prompt, the retrieved context, the model version and the response. Not optional.
We write the choice and the reasoning into handover documentation, because the next team deserves to know why rather than guess - one of several things included in every project build.
Should you use LangChain or write the code yourself?
Use LangGraph if a workflow must survive a restart or pause for a human, the LangChain ecosystem if you need many loaders, and whatever is fastest for a throwaway proof. Keep it if your team already maintains it. Otherwise write the four steps - embed, search, re-rank, prompt - yourself, with tracing. Answer the questions below in order and stop at the first yes.
- Does a workflow need to survive a process restart, or pause for a human for hours? Use LangGraph.
- Do you need more than about ten different document loaders or connectors? Use the ecosystem.
- Are you proving whether an approach works at all, this week? Use whatever is fastest and expect to rewrite it.
- Does your team already know and maintain it? Keep it.
- Otherwise: write the four steps yourself and keep the tracing.
How this plays out in real builds is on RAG systems and AI agent development.
Frequently asked questions
No. A retrieval pipeline is an embedding call, a vector search, a re-rank and a prompt - four steps you can write directly in a couple of hundred lines you fully understand. LangChain helps when you want its ecosystem of loaders and integrations; it does not make the core pipeline meaningfully shorter.
Workflows with real state: branching, loops, checkpoints and human-in-the-loop pauses that survive a process restart. That is a genuine problem and LangGraph solves it well. If your agent is a loop that runs for thirty seconds and either succeeds or fails, you do not have that problem yet.
Debugging distance. When something goes wrong in production, the question is always 'what exactly was sent to the model', and every layer between your code and that HTTP call is a layer you have to see through first. With good tracing this is manageable; without it, it is where the hours go.
Negligibly. Framework overhead is milliseconds against model calls measured in seconds. Choose on maintainability and debuggability, not on performance - performance is not the axis that differs.
Plain application code by default, with the provider SDKs directly. LangGraph when a workflow genuinely needs durable state across hours or days, and Temporal when the durability requirement is stronger than the LLM requirement. We pick per project and we write down why in the handover documentation.
Not at all, and we will not rewrite it for its own sake. Working code with tests is worth more than a preference. We look at whether the abstraction is helping or hiding, and usually the answer is that parts of it earn their place and parts should be flattened.
Inherited a codebase you cannot debug?
Send us the repository. We will tell you which parts of the abstraction are earning their place and which are hiding the problem - on a free call.
