Durable AI Agent Frameworks Compared: One Run, Four Stacks

TL;DR
One agent run crashes mid tool call, waits two days for approval, then retries a charge. How LangGraph, Mastra, Pydantic AI and OpenAI Agents SDK handle it.
The best durable AI agent framework is the one whose state model matches your stack. LangGraph and Mastra checkpoint into your own database and re-execute the interrupted step, while Pydantic AI and the OpenAI Agents SDK get real crash recovery by handing the run to an engine like Temporal or DBOS.
Last updated: October 9, 2026
That one sentence hides the thing that actually bites in production. "Durable" means four different mechanisms across these four stacks, and they disagree on the only question that matters when money or email is involved: after a crash, does your tool call run again?
So instead of comparing feature lists, this page pushes the same failing agent run through each framework and records what happens. I also ran the LangGraph version locally against a SQLite checkpointer so you can see the double execution in real output, not a diagram.
The short answer#
- Python, you already run Postgres, you want explicit control: LangGraph with
PostgresSaveranddurability="sync". Budget for idempotency keys on every side-effecting node, because the interrupted node re-runs from the top. - TypeScript product team: Mastra workflows for deterministic pipelines, Mastra durable agents (still beta) for long agent loops, and the Inngest runner when you need per-tool-call memoization in production.
- Python, you want replay semantics without writing a workflow engine: Pydantic AI plus
DBOSDurability(Postgres or SQLite only) orTemporalDurability(needs a Temporal server). This is the most complete durability story of the four. - You are already on the OpenAI Agents SDK: use
RunStatefor approvals, and add the Temporal integration (GA for Python since March 23, 2026) or the DBOS integration when a crash mid-run must not restart the turn.
If you are still choosing between Mastra and LangGraph on general merit rather than durability, read Mastra vs LangGraph.js first. This page only covers the durability intent.
The test run#
Every framework gets the same refund agent, with three nasty properties:
- Crash mid tool call. Step 2 calls
issue_refund, a non-idempotent side effect. The process dies after the refund API accepted the request but before the framework recorded the result. - Human approval after two days. Step 3 asks a manager to approve a goodwill credit. The manager answers 48 hours later, and you deployed twice in between.
- Retry of a non-idempotent step. When the run comes back, something has to decide whether
issue_refundruns again.
Two mechanisms show up across the four stacks, and naming them makes everything else easier to read:
- Checkpoint and re-execute. The framework saves state at step boundaries. On resume it loads the last checkpoint and re-runs the step that was in flight from its first line. Completed steps are skipped. LangGraph and Mastra work this way.
- Journal and replay. An engine records every completed activity's result in an event history. On recovery it replays your orchestration code and feeds back recorded results instead of calling anything again, then re-executes only the activity that never finished. Temporal works this way, and DBOS checkpoints each completed step so a workflow resumes from the last one that finished.
Neither mechanism can make step 2 safe on its own. The in-flight call re-runs in both models. The difference is how much else re-runs around it.
Side-by-side#
| LangGraph | Mastra | Pydantic AI | OpenAI Agents SDK | |
|---|---|---|---|---|
| Language | Python and JS | TypeScript | Python | Python and TypeScript |
| Where state lives | Checkpointer: SQLite, Postgres, MongoDB, Cosmos DB, or Agent Server | Mastra storage: libSQL by default, Postgres, MongoDB, Upstash, D1, DynamoDB and more | The durable engine you pick: Temporal server, or Postgres/SQLite for DBOS | Nothing mid-run by default; RunState you serialize at an interruption |
| Crash mid tool call | Re-run the interrupted node from the start; completed nodes skipped | Workflows: restart() from last active step. Durable agents: re-run the loop from last snapshot, only with recovery.durableAgents: 'auto' | Engine replays; the unfinished activity or step re-executes | Turn is lost unless you run it on Temporal, DBOS, Restate or Dapr |
| Two-day approval | interrupt() + Command(resume=...); node re-runs from its top on resume | suspend() + resume() with a typed resumeSchema, snapshot in storage | Deferred tools: run ends with DeferredToolRequests, a new run continues with DeferredToolResults | needs_approval tools, result.to_state(), state.approve(), resume |
| Automatic re-drive after crash | Not in the open-source library; your code calls invoke(None, config) | Local server restarts active workflow runs on boot; durable agents opt in | Yes, the engine owns it | Only with an engine integration |
| Extra infrastructure | None beyond your database | None for in-process; Inngest for memoized production runs | Temporal server, or just Postgres/SQLite with DBOS | Temporal server, or Postgres/SQLite with DBOS |
Sources for every cell are linked in the walkthroughs below and listed under Sources.
LangGraph: checkpoints, durability modes, and the node that runs twice#
LangGraph saves "a snapshot of graph state at each super-step," organized into threads keyed by thread_id (checkpointers docs). Production checkpointers ship as separate packages for SQLite, Postgres and MongoDB, plus a Cosmos DB implementation from Azure. If you deploy on LangChain's Agent Server, persistence is handled for you (persistence docs).
Three details decide how the test run goes:
- Durability modes. You pass
durability="exit","async"or"sync"to any execution method.exitonly persists when the graph exits, so the docs state plainly that you "cannot recover from system failures (like process crashes) mid-execution."asyncwrites while the next step runs and carries "a small risk" of a missing checkpoint on crash.syncwrites every checkpoint before the next step starts. For the test run, onlysyncis honest. - Pending writes. If one node in a super-step fails while parallel siblings succeed, the successful writes are kept and those nodes are not re-run on resume.
- Interrupts restart the node. The interrupts docs are explicit: on resume "the runtime restarts the entire node from the beginning," so code before
interrupt()runs again, and side effects placed there "should (ideally) be idempotent."
I wanted to see that rather than trust it. Here is the core of the script I ran on langgraph==1.2.14 with langgraph-checkpoint-sqlite, Python 3.11. Each node bumps a counter in a JSON file outside the graph, which stands in for the outside world:
def charge(state: State):
n = bump("charge_calls") # non-idempotent side effect
if os.environ.get("CRASH") == "1" and n == 1:
os._exit(1) # simulate the process dying mid tool call
return {"charged": True}
def approve(state: State):
bump("approve_pre_interrupt_calls") # code BEFORE interrupt()
decision = interrupt({"question": "Approve refund?"})
bump("approve_post_interrupt_calls")
return {"approved": bool(decision)}
graph = builder.compile(checkpointer=SqliteSaver(sqlite3.connect("checkpoints.db")))
config = {"configurable": {"thread_id": "refund-42"}}
# phase "start": graph.invoke({}, config, durability="sync")
# phase "recover": graph.invoke(None, config, durability="sync")
# phase "approve": graph.invoke(Command(resume=True), config, durability="sync")
Each phase ran as a separate Python process, so nothing survived in memory. The output:
== start (crash)
exit=1
{"plan_calls": 1, "charge_calls": 1}
== recover
next before recover: ('charge',)
interrupt: [Interrupt(value={'question': 'Approve refund?'}, ...)]
ledger: {"plan_calls": 1, "charge_calls": 2, "approve_pre_interrupt_calls": 1}
== approve
final: {'plan': 'refund order 42', 'charged': True, 'approved': True}
ledger: {"plan_calls": 1, "charge_calls": 2, "approve_pre_interrupt_calls": 2, "approve_post_interrupt_calls": 1}
Read the ledger line by line. The completed plan node ran once: the checkpoint did its job. The crashed charge node ran twice: in the real world that is a double refund. The code above interrupt() also ran twice, once when the approval was requested and again when the manager answered. And recovery did not happen on its own: the new process had to call invoke(None, config) after reading next == ('charge',) from get_state().
The fix is not a LangGraph setting. It is an idempotency key derived from something stable across retries. I reran the same crash with the charge keyed on f"{thread_id}:charge" and a dedupe check, and the ledger ended with "charge_node_runs": 2 but exactly one entry under "charges". Payment APIs such as Stripe accept an idempotency key for exactly this reason; pass yours through.
Recent dated changes: langgraph==1.2.12 (September 21, 2026) added response_schema to interrupt(), so the approval payload in step 3 can be validated as a Pydantic model on resume. langgraph==1.2.13 (October 5, 2026) stopped get_state from reporting already-answered interrupts and fixed several DeltaChannel checkpoint paths.
Mastra: workflow snapshots, and durable agents that are still beta#
Mastra has two durability surfaces, and they behave differently.
Workflows. When a step calls suspend(), Mastra persists a snapshot to the workflow_snapshots table in your configured storage, keyed by run ID. libSQL is the default; Postgres, MongoDB, Upstash, Cloudflare D1, DynamoDB and OracleDB are supported (snapshots docs). Resuming is run.resume({ step, resumeData }), and the step's resumeSchema types the manager's answer (suspend and resume docs). For a cold resume after two days and two deploys, you load the run with getWorkflowRunById() and createWorkflowStateReader(), which exposes the suspended step without parsing raw snapshot JSON. Mastra's own durable agents explainer (June 30, 2026) reports a suspended snapshot serializing to 1,059 bytes of JSON and resuming from a fresh process that shared only the database file.
For crashes, restart() resumes "from the last active step," and the local Mastra server restarts all active workflow runs on boot (workflows overview). That is the same checkpoint-and-re-execute shape as LangGraph: the step that was running when the process died runs again. The docs add an autoRestartActiveRuns: false option for workflows "with side effects that must not be re-driven by a blanket restart," which is Mastra telling you the same thing my LangGraph ledger showed.
Durable agents. createDurableAgent() wraps a regular agent so "the agentic loop runs inside a workflow." It was added in @mastra/core@1.45.0 and is labelled beta (durable agents docs). The crash semantics are the part to read twice:
- With the default
recovery.durableAgents: 'off', durable agents persist snapshots only forpending,pausedandsuspendedruns. Approvals survive. In-flight runs do not. - With
'auto', Mastra writes per-steprunningcheckpoints and re-drives orphaned runs on boot. The docs warn this "can reissue LLM calls (with real cost) and re-execute tool calls," and a delegated sub-agent run starts over entirely. - There is no distributed lease yet, so multiple replicas with
'auto'"will race to recover the same runs."
The automatic re-drive is recent. A team running about 28 production agents filed issue #19056 on July 7, 2026 because every deploy silently killed in-flight runs; the fix landed in PR #19191 and the issue closed on July 12, 2026.
For production, Mastra's answer to per-tool-call memoization is createInngestAgent(), where "each tool call becomes a memoized step that Inngest can retry independently." A Temporal runner also exists, announced on May 12, 2026 (Mastra blog) and expanded on May 27, but the workflow runners docs say @mastra/temporal "is experimental and not ready for production use." For more on where Mastra fits beyond durability, see Mastra for durable TypeScript agents.
Pydantic AI: pick an engine, get replay#
Pydantic AI does not ship its own checkpoint store. It supports "eight durable execution solutions, plus a builder for any other engine": Temporal, DBOS, Prefect, Restate and AWS Lambda co-maintained with the vendors, plus Kitaru, Apache Airflow and Absurd (durable execution overview). You attach one as a capability:
agent = Agent(
'openai:gpt-5.6-sol',
instructions="You're an expert in geography.",
name='geography',
capabilities=[TemporalDurability()],
)
That snippet is copied from the Temporal integration docs. Inside a Temporal workflow, model requests, tool calls and MCP communication run as activities; outside one, the same agent behaves normally. The older TemporalAgent wrapper is now marked deprecated.
How the test run goes on Temporal:
- Crash mid tool call. Temporal "relies primarily on a replay mechanism." Completed activities are replayed from history and not called again. The activity that was in flight is "restarted from the beginning," so
issue_refundstill needs an idempotency key. Agentnameand toolsetidvalues become activity names and "should not be changed once the durable agent has been deployed," or in-flight workflows break. - Two-day approval. Pydantic AI's approval path is deferred tools: mark a tool
requires_approval=Trueor raiseApprovalRequired, and the run ends with aDeferredToolRequestsoutput. When the manager answers, you start a new run with the storedmessage_historyand aDeferredToolResultsobject. Nothing is held open for 48 hours, which is why deploys in between do not matter. - Retries. Watch for retry multiplication. The docs note Temporal's default
RetryPolicy.maximum_attemptsof0means unbounded re-execution of model requests, and recommend turning off your provider client's own retries, for examplemax_retries=0on a custom OpenAI client. - Limits. Every payload lands in Temporal's event history, capped at 2MB by default. Base64 encoding means binary content gets roughly 1.5MB of that.
DBOS is the lighter option: the DBOS integration checkpoints model requests and MCP calls "to Postgres or SQLite so a workflow resumes from its last completed step." No separate server to run.
Dated changes from the Pydantic AI releases: v2.50.0 (September 24, 2026) added RunContext.in_durable_context so hooks know they run inside durable workflow code. v2.52.0 (September 29) added workspaces with durable execution support. v2.53.0 (October 1) added AbsurdDurability and made Logfire Temporal spans replay-safe. v2.54.0 (October 2) refuses a second durable engine on the same agent, re-raises model errors from Temporal activities with their original type, and adds RunCancelled.from_cancellation() for cancelled DBOS workflows.
OpenAI Agents SDK: great approvals, durability by integration#
The SDK's human-in-the-loop flow is the cleanest approval API of the four. Mark a tool with needs_approval=True (or a callable that decides per call), and the run pauses with ToolApprovalItem entries in result.interruptions. Convert with result.to_state(), call state.approve(...) or state.reject(...), and resume with Runner.run(agent, state). The docs say RunState "is designed to be durable": serialize it with to_json(), store it, and rebuild it two days later with RunState.from_json().
Two warnings from the same page matter for step 3. RunState.from_json() does not authenticate the snapshot, so keep it in server-side storage and never accept serialized state back from a browser. And if approvals may sit for a while, store a version marker for your agent definitions next to the state, because prompts and tools can change before the manager clicks approve.
What the SDK does not do on its own is survive step 1. RunState exists at interruptions, not between every tool call, so a crash mid issue_refund loses the turn. For that, the running agents docs list four integrations: Temporal, Restate, DBOS ("requires only a SQLite or Postgres database") and Dapr.
The Temporal one is the most mature. It entered public preview on July 30, 2025 and, per Temporal's update, became generally available for the Python SDK on March 23, 2026. Temporal provides its own Runner implementation so that "every agent invocation is executed through a Temporal Activity," and tools become activities via activity_as_tool (Temporal cookbook). Temporal also documents a TypeScript package, @temporalio/openai-agents, but the GA announcement covers Python. For how this SDK compares with Anthropic's on everything else, see OpenAI Agents SDK vs Claude Agent SDK.
Where Vercel WorkflowAgent fits#
TypeScript teams on the AI SDK have a fifth option with a dated primary source. AI SDK 7, released June 25, 2026, introduced WorkflowAgent in @ai-sdk/workflow. Vercel's changelog says execution state "is persisted to durable storage between steps, so agents survive deploys, process restarts, interruptions, and delayed approvals." It replaces the Workflow DevKit's DurableAgent, makes approval a needsApproval property on the tool, and runs each tool call as a durable step (WorkflowAgent docs). It is the most direct path if you deploy on Vercel already; the programming model is covered in Vercel's durable execution model.
How to choose#
Ask three questions in order.
- What language is the product in? Mastra and WorkflowAgent are TypeScript-only. Pydantic AI is Python-only. LangGraph and the OpenAI Agents SDK cover both, but the durability integrations are deepest in Python.
- Are you willing to run an engine? If no, your realistic options are LangGraph checkpointers, Mastra storage, or DBOS on Postgres or SQLite. If yes, Temporal gives you replay semantics and automatic re-drive under either Pydantic AI or the OpenAI Agents SDK.
- How many side effects sit inside a single step? Checkpoint-and-re-execute is fine when every step does one thing and is keyed for idempotency. When one agent turn makes many tool calls, you want per-tool-call memoization: Temporal activities, DBOS steps, Inngest steps under Mastra, or WorkflowAgent steps.
My default for a new Python agent that touches money is Pydantic AI on DBOS: replay semantics, no new server, and state in the Postgres you already back up. For a TypeScript product, it is Mastra workflows with explicit suspend() points for approvals, and durable agents only once the beta label comes off. Whatever you pick, the idempotency key is not optional, and treating a long run as something a harness has to supervise beats trusting the framework to do it.
What people are actually saying#
- Practitioners hit the deploy problem first, not the crash problem. The author of Mastra issue #19056 runs about 28 production agents on a self-hosted Mastra server and reported that every deploy killed in-flight runs silently. Their pushback was specific: switching to Inngest meant "adopting a new vendor/infra just to survive our own deploys" when the snapshot already lived in their Postgres. Mastra shipped recovery five days later.
- The idempotency question never goes away. In the January 2025 DBOS TypeScript launch thread on Hacker News (77 points), one commenter asked whether durability includes idempotency when step 2 calls a payment API after step 1 wrote a row. The answer from DBOS's Peter Kraft, the Show HN author, was a saga: durability guarantees the compensating steps run, not that external calls happen once.
- The sharpest disagreement is Postgres versus a dedicated engine. In the same thread, a commenter who disclosed being a former Temporal employee questioned what happens when the Postgres server fills up under load. DBOS replied that it scales as far as Postgres does and cited 10K+ steps per second on a large server. On code changes mid-run, DBOS was blunt: "every workflow finishes on the code version it started," which is the same constraint behind Pydantic AI's rule that activity names never change after deploy.
- The loudest "checkpoints are not durable execution" takes come from engine vendors. Diagrid's February 2026 post and Restate's June 2026 post both argue that framework checkpointing leaves failure detection, resumption, retries and idempotency to you. My ledger output agrees on the facts. Both HN submissions drew fewer than five points, though, and the counter-case is real: if every side effect is keyed and a sweeper re-drives stuck threads, a checkpointer plus your existing database covers most teams without a new system to operate.
FAQ#
What is a durable AI agent framework?#
A durable AI agent framework persists an agent run's progress outside the process, so the run can resume after a crash, a deploy, or a long human wait without starting over. In practice that means checkpoints in a database you configure (LangGraph, Mastra) or an event history kept by a durable execution engine such as Temporal or DBOS (Pydantic AI, OpenAI Agents SDK integrations).
Does LangGraph re-run a tool call after a crash?#
Yes. The node that was executing when the process died runs again from its first line when you resume the thread, and code before an interrupt() also runs again when an approval is answered. I confirmed both on langgraph==1.2.14 with a SQLite checkpointer. Completed nodes are not re-run. Use durability="sync" and an idempotency key for any side effect.
Do I need Temporal for durable agents?#
No. LangGraph needs only a checkpointer database, Mastra needs only its configured storage, and DBOS works with Postgres or SQLite through both Pydantic AI and the OpenAI Agents SDK. Temporal is worth running when you want replay semantics, automatic re-drive, per-activity retry policies and an execution UI, and you are willing to operate a server.
Which durable agent framework works in TypeScript?#
Mastra and Vercel's WorkflowAgent from AI SDK 7 are TypeScript-first. LangGraph.js uses the same checkpointer model as the Python library. The OpenAI Agents SDK has a TypeScript version with the same approval flow, and Temporal documents a TypeScript integration, though its March 23, 2026 GA announcement covers Python.
How long can an agent wait for human approval?#
As long as your storage keeps the state. LangGraph says the graph "waits indefinitely" after interrupt(), Mastra snapshots persist across deployments and restarts, Pydantic AI's deferred tools end the run so nothing is held open, and OpenAI's RunState can be stored as JSON. Store a version marker with the state, because your code may change before the answer arrives.
Continue Reading#
- Mastra vs LangGraph.js - the general framework comparison beyond durability
- Mastra for Durable TypeScript Agents - where Mastra fits and where it does not
- Vercel's Durable Execution Programming Model - the platform layer under WorkflowAgent
- OpenAI Agents SDK vs Claude Agent SDK - the two big platform SDKs side by side
- Long-Running Agents Need Harnesses - supervision, stop conditions and receipts for long runs
Sources#
- LangGraph persistence docs - checkpointers, stores, Agent Server persistence
- LangGraph checkpointers docs - super-steps, pending writes, durability modes, checkpointer packages
- LangGraph interrupts docs - resume semantics and idempotency rules
- LangGraph 1.2.12 release and LangGraph releases - September 21 and October 5, 2026 changes
- Mastra suspend and resume, snapshots and workflows overview
- Mastra durable agents docs - beta status, recovery modes, multi-replica caveat
- Mastra workflow runners, Introducing Temporal Support for Mastra Workflows (May 12, 2026) and Mastra Workflows, Enhanced (May 27, 2026)
- Mastra: What are durable AI agents? (June 30, 2026)
- Mastra issue #19056 - orphaned running runs, closed July 12, 2026
- Pydantic AI durable execution overview, Temporal integration, DBOS integration and deferred tools
- Pydantic AI releases and v2.54.0
- OpenAI Agents SDK docs, human-in-the-loop and running agents
- Temporal: OpenAI Agents SDK integration announcement and GA update and Temporal AI cookbook
- Vercel: AI SDK 7, AI SDK 7 changelog and WorkflowAgent docs
- Stripe idempotent requests
- Community: DBOS TypeScript on Hacker News, Diagrid post on Hacker News, Restate post on Hacker News
Get the next comparison like this in your inbox
One email a week on AI Agents and the rest of the AI dev stack. Free.
Read next on AI coding tools
Mastra vs LangGraph.js: TypeScript Agent Frameworks Head to Head
Both Mastra and LangGraph.js are serious TypeScript agent frameworks - but they start from opposite philosophies. Here is what that means for your next project.
8 min readVercel's New Durable Execution Programming Model: A Developer's Guide
Durable execution lands on Vercel. What it means for agents, long-running flows, and indie dev stacks - with code, gotchas, and where it fits the agent stack.
10 min readOpenAI Agents SDK vs Claude Agent SDK: Building Agents on the Two Big Platforms
A practical comparison of OpenAI's Agents SDK and Anthropic's Claude Agent SDK - orchestration models, tool ecosystems, sandboxing, and how to choose the right platform for your team.
9 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.







