Your agent cannot start itself, survive a hangup, or survive a crash, and those are three different bugs
'Nobody is holding the connection' is three unrelated failures with three separate opt-ins, and teams keep reaching for the wrong one.

Foundry routines, background mode, and resilient execution each fix exactly one of them and are silent about the other two. This article names the one you actually need.
It is 7 AM, nobody is logged in, and the three-day review died on day two. Both are the same missing feature, and neither one is a bigger model.
In this article: You will learn why "nobody is holding the connection" is three unrelated failures, how Foundry routines collapse a scheduler, a Function, and a pile of authentication code into one project resource, why a scheduled agent's dispatch identity is fixed at creation time, and how crash recovery reenters your handler from the top. By the end you will know which flag answers which failure, and which half of the job never leaves your code.
Picture a research desk: one agent, three hundred research analysts, strict per-analyst isolation, and a clean audit story. It works. It also cannot do three things.
Nothing runs at seven in the morning, before the analysts arrive. A multi-hour deep review of a tier-one covered company holds an HTTP connection open and hopes. And when the container stops mid-review, the work in flight is gone even though every finding written so far is durable.
Those three failures sound like one problem. They are not. They have three different owners, three different opt-ins, and three different documentation pages, and teams routinely reach for the wrong one. A scheduler does not save a crashed run. Crash recovery does not make an agent start itself. Sorting that out is most of the work here.
One caution first. Routines are documented only on Microsoft Learn, with no independent corroboration available, so every claim below is mirror-grounded and the concept page is dated August 27, 2026 [PERISHABLE: checked September 2026]. Do not import assumptions from another platform's scheduler. This one has its own identity rule, and it is not the one you would guess.
Three answers to "nobody is holding the connection"
Work that outlives a request fails in three distinct places, and Foundry gives you a separate mechanism for each.

Routines are the trigger layer: a project-native rule that says when a time, a schedule, or an event occurs, invoke this agent. Without them, teams build that layer out of schedulers, Logic Apps, Azure Functions, queues, storage, and authentication code. Routines move the glue into Foundry, so the trigger, the action, the permissions, the connections, and the run history live with the agent.
Background mode and resilient execution solve different problems, and the docs draw the line in one sentence worth memorizing: background mode lets work continue after the initiating request returns, and resilient execution preserves work after the hosting process stops. The resilience concept page splits that into three capabilities, each with an explicit column for what it does not give you.
| Capability | What it provides | What it does not provide |
|---|---|---|
| Background execution | Asynchronous work that clients can poll or reconnect to. | Recovery after the process that owns the work stops. |
| Resilient execution | Durable work identity, persisted input, process-loss detection, and handler reentry. | Automatic preservation of every intermediate application state or side effect. |
| Stream replay | Retained events that reconnecting clients can receive from a cursor. | A checkpoint of the agent's internal workflow state. |
Each capability leaves one gap:
- Background execution: asynchronous work that clients can poll or reconnect to.
- Background execution does not provide: recovery after the process that owns the work stops.
- Resilient execution: durable work identity, persisted input, process-loss detection, and handler reentry.
- Resilient execution does not provide: automatic preservation of every intermediate application state or side effect.
- Stream replay: retained events that reconnecting clients can receive from a cursor.
- Stream replay does not provide: a checkpoint of the agent's internal workflow state.
Read the right-hand column twice. Each row solves exactly one failure and is silent about the other two, which is why "we turned on background mode" is not an answer to "the container restarted."
Long-running agent resilience is in preview, and its APIs and package versions carry the usual preview caveat [PERISHABLE: checked September 2026]. Routines are not marked preview. The reminder tool inside them is.
Foundry routines: one trigger, one action, and no scheduler to run
A routine is deliberately not an orchestrator, and that constraint is what makes it cheap.
A routine holds five things: a trigger, an action that invokes exactly one prompt agent or hosted agent, an input sent as text or JSON, a lifecycle state you can flip without touching the agent, and a run history with each run's inputs, outputs, status, and a link to the trace. One trigger, one action. The docs state the whole model as a single question: when should this agent run?

The concept page names three trigger types, timer, recurring, and event. The API spells the recurring one schedule and splits the event case into two concrete triggers.
| Trigger type | Fires | Required fields |
|---|---|---|
timer |
Once, at a specific future date and time. The routine then becomes inactive. | at, an ISO 8601 timestamp with an explicit UTC offset |
schedule |
Repeatedly, on a five-field cron expression, at a minimum interval of five minutes | cron_expression and time_zone |
github_issue |
When an issue is opened or closed in a watched repository | connection_id, owner, repository, issue_event |
custom with the teams provider |
When a message is posted to a watched Teams channel | provider, event_name, and a parameters object |
The four trigger types differ in when each fires and what each one requires:
timer: fires once, at a specific future date and time, after which the routine becomes inactive. Required fields:at, an ISO 8601 timestamp with an explicit UTC offset.schedule: fires repeatedly on a five-field cron expression, at a minimum interval of five minutes. Required fields:cron_expressionandtime_zone.github_issue: fires when an issue is opened or closed in a watched repository. Required fields:connection_id,owner,repository,issue_event.customwith theteamsprovider: fires when a message is posted to a watched Teams channel. Required fields:provider,event_name, and aparametersobject.
The desk's weekday-morning sweep is the part a research team would actually notice. The Python SDK reaches routines under the beta namespace, and it needs azure-ai-projects 2.4.0 or later.
The whole routine is a single upsert call carrying two nested objects, one describing when the agent runs and one describing what gets invoked. Nothing outside those two objects is a runtime concern.
routine = client.beta.routines.create_or_update( # ①
routine_name="weekday-topic-sweep",
description="Sweeps tier-one covered companies before the desk opens.",
enabled=True, # ②
triggers={
"weekday-morning": {
"type": "schedule", # ③
"cron_expression": "0 7 * * 1-5", # required ④
"time_zone": "America/Chicago", # required ⑤
}
},
action={
"type": "invoke_agent_responses_api", # ⑥
"agent_name": "research-desk", # required ⑦
"input": "Sweep tier-one covered companies for filings and news since yesterday.", # ⑧
},
)
① create_or_update is an upsert keyed by routine_name, so rerunning the script edits the existing routine rather than creating a second one.
② The lifecycle state is a field on the routine, which is how you stop a schedule without touching or redeploying the agent it invokes.
③ The trigger type selects the schema for the rest of the object. A schedule trigger fires repeatedly on a cron expression, at a minimum interval of five minutes.
④ The cron expression is five fields, here 7 AM on Monday through Friday.
⑤ time_zone is required rather than optional, and it takes an IANA or Windows identifier. A routine created through the API pins the zone explicitly instead of inheriting a browser's local zone.
⑥ The action type has to match the protocol the agent actually exposes. The desk declared Responses, so invoke_agent_responses_api is the correct one.
⑦ Name the agent with agent_name or agent_endpoint_id, never both.
⑧ The input is what Foundry sends the agent on dispatch, as text or JSON.
Note: The full extracted listing at code/azure-foundry-hyperscaler/part-10-work-that-outlives-the-request/listings/01-weekday-topic-sweep.py shows the client construction elided here.
One detail the callouts understate: a routine created in the portal inherits the browser's local zone, while one created through the API pins the zone explicitly. Use invoke_agent_invocations_api instead when the hosted agent speaks Invocations.
Nothing about a routine introduces a second runtime. Foundry creates a run record and invokes the agent endpoint with the routine input, and the agent answers with the same model, instructions, tools, and observability it uses when a human calls it. Because the dispatch reuses the existing invocation path, routines work against virtual-network-secured projects with no extra networking setup.
Gotcha: routines do not support customer-managed key encryption, they do not support workflow agents, and they are unavailable in UK West, Switzerland West, Japan West, UAE North, and Norway East [PERISHABLE: checked September 2026]. If Routines is missing from the portal navigation, the feature is not enabled for that region or subscription rather than hidden behind a permission. Creating and managing routines needs Foundry User or higher on the project.
The dispatch identity is fixed at creation, and the default will bite you
Every routine runs as the agent unless you opt in at create time, the opt-in means one specific human, and changing your mind means deleting the routine.
A routine has no user token by construction, because no user is there. The docs state the consequence in bold and then repeat it in the limitations list: every routine invokes its agent using agent identity by default, regardless of trigger type. The agent's own role assignments, the ones that belong on agentIdentityId rather than on the project managed identity, determine everything the run can reach.

When the agent has tools that require delegated user access, you opt in to creator identity with one top-level object on the create request:
{
"authorization": { "identity": "creator" },
"triggers": { "...": { "type": "..." } },
"action": {
"type": "invoke_agent_responses_api",
"agent_name": "<your-agent-name>"
}
}
Creator identity means the Microsoft Entra identity of the principal that created the routine, and the docs enumerate everyone it is not: not the agent creator, not the agent publisher, not the connection creator, not a later routine editor, and not another end user. You cannot supply an arbitrary end-user identity at dispatch time, and if the creator's access or consent is later removed, those tool calls fail.
Gotcha: the authorization object is accepted only on create. An update ignores it silently, so switching an existing routine between agent and creator identity means deleting and recreating it. Two traps sit next to that one. The current SDK and azd routine models use the default agent identity, so creator identity is a REST-only gesture today [PERISHABLE: checked September 2026]. And connector authentication is separate from dispatch identity, so it does not change the identity stored in a GitHub or Teams connection.
For the desk this is a design decision, not a flag. Archive reads run as the signed-in analyst through OAuth identity passthrough, and a 7 AM sweep has no signed-in analyst. Creator identity would run the whole sweep as one research service principal, which needs organization-wide archive access to be useful and therefore collapses the per-analyst entitlement model. The desk does the other thing: the sweep touches only public filings and news, and it writes per-topic findings that each analyst's first interactive turn cross-checks against their own archive slice under their own token.
Scheduled work runs unattended and stays on unattended-safe tools. The rule this feature teaches is a general one.
Events, run history, and why a completed run is not a finished job
The event triggers are real integrations with their own identity, and the run record tells you about delivery rather than about outcome.
Two event triggers ship. github_issue fires when an issue is opened or closed in a watched repository, and custom with the teams provider and event name on_new_channel_message fires when a message is posted to a watched channel. Both authenticate through a connector connection Foundry provisions in your account's connector namespace, referenced from the trigger by connection_id.
One behavior surprises people on the first test run. For both event triggers the event payload overwrites action.input when the event fires, so the input you configured applies only to manual test dispatches. Your carefully worded prompt is not what the agent receives in production, and instructions written to assume a prompt rather than a payload are the resulting bug.
Run history answers what an on-call engineer actually asks: did the trigger fire, what input went to the agent, did the invocation complete or fail, and which trace holds the model, tool, and latency detail. In Python it is one call, and the failure fields are on the run object:
for run in client.beta.routines.list_runs("weekday-topic-sweep"):
print(run.id, run.phase, run.attempt_source, run.started_at, run.ended_at)
if run.phase == "failed":
print(run.error_type, run.error_message)
Test a routine before its schedule arrives by dispatching it manually with :dispatch_async, the only dispatch route in the public contract, which returns a dispatch_id and a task_id.
Gotcha: acknowledgment is not completion, twice over. A :dispatch_async response means the run was enqueued, and a completed run means the downstream API returned success for the dispatch request, not that asynchronous work the agent started has finished. The delivery mechanics are worth knowing before you debug one: three total attempts with exponential backoff starting at one second and capped at five, a 30-second per-attempt timeout on the downstream HTTP call that excludes queueing and backoff, and 408, 429, and 5xx treated as retryable while other 4xx are terminal. Internalize that 30-second ceiling: a routine cannot wait for a multi-hour review, which is exactly why the rest of this article exists.
A routine is a trigger, not an orchestrator. The moment the automation needs branching, multiple agents, human approval, or complex state, the docs send you to a workflow instead, and the comparison draws the line as trigger-to-agent against a graph of nodes, edges, branching, and state.
The reminder tool: when the agent schedules itself
Routines are external triggers. The inverse case is an agent deciding mid-run that it needs to come back later.
A hosted agent can schedule itself with the built-in reminder_preview tool, exposed through a toolbox. The agent supplies a delay in minutes and instructions for its future self, and Foundry re-invokes the same agent when the delay elapses, implementing the delay by creating a scheduled routine.
The tool is connectionless, so it needs no credential and no external service, and it drops into an existing toolbox file as one more entry:
tools:
- type: reminder_preview
name: schedule_reminder
description: Schedule a reminder that re-invokes this agent at a future time.
Two arguments, and the ranges matter: minutes accepts 1 through 43,200, which is thirty days, and input carries the instructions for the follow-up run. The model decides when to call it based on your instructions, and you write no code to handle the reminder invocation.
The difference from a routine is worth carrying: a reminder re-invokes the agent on the same conversation, preserving context, while an ordinary routine starts a new one unless its action names an existing conversation. A reminder is how a review waiting on a quarterly filing checks back on its own, without a human queueing it and without the container staying alive in between.
Gotcha: the reminder tool is hosted-agents only and is in preview [PERISHABLE: checked September 2026]. A prompt agent cannot use it, which means a team that prototyped in the playground and plans to "add reminders later" has a rewrite ahead of them rather than a configuration change.
Background mode: the request returns and the work keeps going
One request flag, platform-managed polling, and a boundary that catches people the first time a container restarts.
On the Responses protocol, background mode is the documented answer for long-running tasks: set background to true and poll the response status until it completes. The protocol comparison lists it as one of the reasons to choose Responses at all, with platform-managed polling and cancellation.
The listing is two halves that no longer have to run in the same place: a submission that returns immediately, and a poll loop that any later client can run against the returned ID.
response = openai.responses.create(
extra_body={"agent_reference": {"name": AGENT_NAME, "type": "agent_reference"}}, # ①
input="Run the deep review for Aldrich Components.",
background=True, # ②
)
while response.status in ("queued", "in_progress"): # ③
sleep(2)
response = openai.responses.retrieve(response.id) # ④
① The agent reference rides in extra_body, which is how a stock OpenAI client names a Foundry agent without a Foundry-specific client type.
② This one flag is the entire opt-in. The call returns as soon as the work is accepted rather than when it finishes.
③ The status values to wait on are queued and in_progress. Polling is platform-managed, so there is no callback to register and no connection to hold.
④ Retrieval is by response ID, which is why the poll loop does not have to be the process that submitted the work.
Note: The full extracted listing at code/azure-foundry-hyperscaler/part-10-work-that-outlives-the-request/listings/02-background-deep-review.py shows the client construction and imports elided here.
The caller can now disconnect, come back, and retrieve the same response by ID. Streaming and background mode change how the application receives a response, not the relationship among agents, conversations, and responses.
Gotcha: background mode survives the caller leaving and does not survive the container leaving. Opting in to recovery is a separate step, and the docs say so on the same line that recommends background mode. Foreground responses, meaning background=false, are always marked failed on a crash, because their client connection is already gone. No configuration changes that.
Resilient execution: a lease, two identities, and a handler you can reenter
Recovery reenters your handler from the top with the same input, so everything that makes recovery useful is durable progress you wrote yourself.
A resilient unit of work has two identities: a work identity naming the logical job or multi-turn conversation, and an input identity naming one input or turn within it. Before a handler starts, the runtime persists its input and takes a lease on the work record, renewing it while the handler runs. If the process stops and abandons the lease, a later process reclaims the record and invokes the registered handler with the same identities and input.

Three sentences from that page prevent most of the misunderstandings. Recovery reenters the handler from its beginning; it is not deterministic replay, and it does not restore local variables or an in-memory call stack. Recovery is not retry: a retry handles a failure the running handler reported and consumes retry budget, while recovery continues the same durable attempt after its process disappears. And the handler uses durable checkpoints or watermarks to work out what is already done.
There are two work shapes. One-shot work is one input producing one result, and calls sharing its work identity converge on the same logical operation instead of duplicating it. A multi-turn chain keeps one work identity active across turns until the application deletes it or retention expires. A weekly brief is one-shot. A three-day review conversation is a chain.
The responsibility split is the honest part.
| Foundry and the AgentServer SDK provide | Your agent provides |
|---|---|
| Durable work and input identities | Stable identifiers that map to your job or conversation model |
| Input persistence before handler execution | Inputs inside the task payload limit, with large data stored externally |
| Lease-based process-loss detection and handler reentry | A safe rerun path or a recovery branch that resumes from durable progress |
| Small durable metadata values | Checkpoint references, idempotency keys, and side-effect watermarks |
| Conversation locking and optional steering queues | User-visible behavior for queued, rejected, interrupted, and canceled turns |
| Event retention and cursor-based replay | Per-turn stream identities and reconnect-aware clients |
| Cleanup of terminal or expired runtime records | Cleanup of external checkpoints, sessions, and application data |
The same split as a list pairs each platform-supplied mechanism with the thing you still have to supply:
- Platform provides durable work and input identities. You provide stable identifiers that map to your job or conversation model.
- Platform provides input persistence before handler execution. You provide inputs inside the task payload limit, with large data stored externally.
- Platform provides lease-based process-loss detection and handler reentry. You provide a safe rerun path or a recovery branch that resumes from durable progress.
- Platform provides small durable metadata values. You provide checkpoint references, idempotency keys, and side-effect watermarks.
- Platform provides conversation locking and optional steering queues. You provide user-visible behavior for queued, rejected, interrupted, and canceled turns.
- Platform provides event retention and cursor-based replay. You provide per-turn stream identities and reconnect-aware clients.
- Platform provides cleanup of terminal or expired runtime records. You provide cleanup of external checkpoints, sessions, and application data.
Keep task metadata small and treat it as a checkpoint index rather than a checkpoint store: a session or checkpoint ID, the last completed phase, an idempotency key, or a pointer to a database or blob. Conversation history, model output, tool results, and large artifacts belong in a framework checkpointer or your own storage, which for the desk is the reviews/{review_id} state store. The task input limit is roughly 10 MiB after JSON serialization, raising InputTooLarge before any network call.
Turning recovery on, and the half you still write
Crash recovery is off by default, one flag buys you the framework half, and resuming instead of rerunning is the half that lands on your state store.
Enable it explicitly for the surface you use. On the Responses protocol it is one option on the host:
from azure.ai.agentserver.responses import ResponsesAgentServerHost, ResponsesServerOptions
app = ResponsesAgentServerHost(
options=ResponsesServerOptions(resilient_background=True),
)
Recovery applies only to responses that are both stored and backgrounded, meaning store=true and background=true. Without the flag, a crashed background response is marked failed with error.code="server_error" and the handler is never reinvoked. With it, you get handler reinvocation with the same request, input, and metadata, stream replay of persisted SSE events, a conversation lock against concurrent writes, and cleanup that marks nonrecoverable responses failed rather than silently rerunning them.
On the task primitives the opt-in is implicit: declaring a @task or @multi_turn_task handler automatically enables the startup recovery scan, and set_resilient_tasks_enabled(True) at import time is the escape hatch for tasks registered lazily after host startup.
A recovered handler that changes nothing still produces a correct response. It simply reruns the whole turn. Resuming is the optional half, and you branch on a marker: context.is_recovery on the Responses side, or ctx.entry_mode on the task side, whose values are fresh, resumed, and recovered.
Four resume strategies are documented, and the desk's choice is the third.
| Strategy | Where progress lives | Recovery behavior |
|---|---|---|
| Naive rerun | Nowhere | Rerun the whole turn. Correct, and unsafe for unfenced non-idempotent side effects. |
| Framework checkpoints | Persisted response snapshots | Seed from context.persisted_response, resume after the checkpointed output items. |
| Upstream-owned resume | Your framework or application store | Rebuild from a framework checkpoint or your own database. |
| Watermark overlay | Small metadata watermarks | Combine with any strategy to avoid repeating side effects. |
The four resume strategies differ in where each one keeps progress and how each behaves on recovery:
- Naive rerun: progress lives nowhere, and recovery reruns the whole turn. Correct, and unsafe for unfenced non-idempotent side effects.
- Framework checkpoints: progress lives in persisted response snapshots, and recovery seeds from
context.persisted_responseand resumes after the checkpointed output items. - Upstream-owned resume: progress lives in your framework or application store, and recovery rebuilds from a framework checkpoint or your own database.
- Watermark overlay: progress lives in small metadata watermarks, and it combines with any strategy to avoid repeating side effects.
The desk's state store already implements upstream-owned resume without calling it that. record_review_step writes each finding under a deterministic findings/{step} key and only then advances a progress marker guarded by an ETag. Reading that marker back is the entire recovery branch.

The listing is a reader and a loop. next_step turns the durable progress marker into an integer, and the handler runs phases from that integer forward, with one branch on entry and one on shutdown.
async def next_step(review_id: str) -> int:
store = await FoundryStateStore.get_or_create( # ①
f"reviews/{review_id}", user_isolation=True
)
async with store:
progress = await store.get_item("progress")
return int(progress.value["workflow_step"]) if progress else 0 # ②
@app.response_handler # ③
async def handler(request, context, _cancellation_signal):
review_id = await review_id_for(request)
step = await next_step(review_id) if context.is_recovery else 0 # ④
while step < len(REVIEW_STEPS):
if context.is_shutting_down: # ⑤
await context.exit_for_recovery() # stays in_progress for a later lifetime ⑥
return
await record_review_step(review_id, step, await REVIEW_STEPS[step](review_id)) # ⑦
step += 1
① The recovery branch reads the same reviews/{review_id} store the desk has always written to. No second persistence story is introduced for resilience.
② A missing marker means step zero, so a first run and a crash before any write take the same path.
③ The handler is registered rather than called. Reentry after a process loss invokes this same function with the same request, input, and metadata.
④ context.is_recovery is the only thing the platform tells you about reentry. Branching on it is what turns a rerun into a resume.
⑤ Shutdown is checked at a phase boundary, which is the point where durable progress and in-memory progress agree.
⑥ exit_for_recovery() defers without writing a terminal state, so the record stays in progress and a later lifetime reclaims it. Returning an error here would end the work instead.
⑦ Each phase writes its finding and advances the marker through the store's write-order protocol, which is what makes the value read at
② trustworthy.
Note: The full extracted listing at
code/azure-foundry-hyperscaler/part-10-work-that-outlives-the-request/listings/03-resumable-review-handler.py
shows the host construction, record_review_step, and the phase functions elided here.
Nothing in those lines is new state machinery. The resilience feature supplies reentry and the marker that tells you it is a reentry. Your write-order protocol supplies the answer to "where was I." Prefer phase boundaries that checkpoint cleanly, completing one output item per phase and then checkpointing, so a crash before a checkpoint reruns one phase and a crash after it skips that phase entirely.
Graceful shutdown is deliberately different from terminal failure, which is why exit_for_recovery() appears in that loop. Crash recovery reenters the same attempt: it does not consume retry budget, and a wall-clock timeout does not reset because the process restarted.
Gotcha: recovery makes side effects happen twice unless you fence them. Before an action an upstream system cannot deduplicate, stamp a watermark and flush it, then clear it after the operation commits, and check it on reentry:
context.conversation_chain_metadata["brief_posted"] = True
await context.conversation_chain_metadata.flush() # fence before the side effect
await post_weekly_brief_to_teams(review_id)
For the desk, that is the difference between a restart costing thirty seconds of recomputation and a restart posting the same weekly research brief to a research channel twice.
Stream replay, so a reconnecting client gets what it missed
The cursor that lets a client catch up is the same cursor that lets a crashed producer resume, and picking the backing is a startup decision you make once.
Two rules govern streaming for long-running work. Use a per-turn stream ID, and never reuse a multi-turn work or conversation ID as the stream ID, because a completed stream closes while the conversation continues. And choose a backing that matches the recovery you need, because the backing decides whether late subscribers can replay and whether the stream survives a restart.
With the Invocations and task primitives you pick one backing at startup and then look streams up by ID anywhere in the process.
| Backing | Replay for late or reconnecting subscribers | Survives process restart |
|---|---|---|
use_in_memory_live(), the default |
No | No |
use_in_memory_replay(...) |
Yes, within ttl_seconds |
No |
use_file_backed_replay(...) |
Yes | Yes |
The three stream backings each guarantee something different:
use_in_memory_live(), the default: no replay for late or reconnecting subscribers, and it does not survive a process restart.use_in_memory_replay(...): replay withinttl_seconds, and it does not survive a process restart.use_file_backed_replay(...): replay for late or reconnecting subscribers, and it survives a process restart.
Pass cursor_fn when you want cursored reconnect, because without it subscribe(after=...) is ignored and last_cursor() returns None. After a crash, a file-backed producer reads the last persisted cursor and continues from the next one, which is why the docs call the cursor both the client's reconnect primitive and the producer's recovery primitive. Do not mirror stream cursors into task metadata; the stream log already owns stream progress.
On the Responses protocol the platform manages the stream and the client reconnects with a query parameter:
GET /responses/{id}?stream=true&starting_after=<sequence_number>
Gotcha: teach your clients that any response.in_progress event after the first one is a snapshot reset. The client replaces its accumulated output with that snapshot, discards partial output missing from it, and applies later events additively. Output indexes are slot identifiers rather than monotonic counters. A client that appends blindly renders a review twice and looks like an agent bug.
One adjacent feature shares these options. Setting steerable_conversations=True on the same ResponsesServerOptions makes a second turn arriving on a busy conversation queue behind the active one and cooperatively cancel it, instead of returning 409 conversation_locked. An analyst redirecting a running review mid-sentence is the case it exists for.
Both frameworks, and one gap the docs leave open
Routines never touch your framework. The trigger lives in the project, the action names an agent and a protocol, and the dispatch goes through the ordinary invocation path, so a LangGraph container and a Microsoft Agent Framework container are scheduled by identical routine definitions. The reminder tool is the same story, arriving as an MCP tool from a toolbox.
Resilience sits one layer lower and is still shared: the task primitives, the streaming registry, and FoundryStateStore all live in azure-ai-agentserver-core, which both tracks install. Version floors move here, so take the higher one. The long-running reference page requires azure-ai-agentserver-core 2.0.0 or later in Python and Azure.AI.AgentServer.Core 1.0.0-beta.28 or later in .NET, above the 2.0.0b7 floor cited elsewhere [CONTESTED: agents-concepts-long-running-agent-reference.md against agents-how-to-isolate-sessions-per-user.md, take the higher floor].
One gap is worth naming rather than papering over. Every sample that sets resilient_background=True constructs ResponsesAgentServerHost from the raw protocol library, and neither hosting page shows a framework bridge accepting server options: the LangGraph page writes ResponsesHostServer(graph).run(port=port) and the Agent Framework page writes ResponsesHostServer(agent), with no options argument anywhere. The capability is in the shared layer and the documented wiring is not. Check the installed bridge's signature before you plan on it, and fall back to the raw host with your framework invoked inside the handler if it is absent.
Do this today
- List the work your agent cannot do today and sort it into the three buckets: nobody started it, the caller hung up, the process died. Name the mechanism for each before you write any code.
- Create one
scheduleroutine with an explicittime_zoneand a description a stranger can read at 3 AM, then dispatch it manually with:dispatch_asyncbefore its first scheduled fire. - Audit the tools your scheduled agent will reach. Anything requiring delegated user access does not work under the default agent identity, and creator identity is a create-time, REST-only decision you cannot change later.
- Turn on
resilient_background=True, then add exactly one thing: acontext.is_recoverybranch that reads a durable progress marker you already write. - Find every non-idempotent side effect the agent performs, such as posting to a channel or filing a ticket, and fence each one with a flushed watermark before shipping recovery.
The desk runs itself
At 7 AM Chicago time on weekdays, weekday-topic-sweep fires. Foundry creates a run record, invokes research-desk under the agent identity, and the sweep pulls overnight filings and news for every tier-one covered company. Findings land in the state store, keyed per topic, and an analyst arriving at 8 AM sees work that was already done. Nobody provisioned a scheduler, wrote authentication code, or ran a Function.
A deep review is the other half. An analyst starts one, the desk creates a stored background response, and the analyst closes the laptop. A platform restart at phase four reenters the handler with context.is_recovery true, next_step reads the marker, and the review resumes at phase four instead of phase zero. The client reconnects with starting_after and gets the events it missed. When the review needs a quarterly filing that will not exist for three weeks, the desk calls schedule_reminder and comes back to the same conversation on its own.
Foundry takes over the trigger layer outright, and it takes over the durable half of long-running work. What stays yours is every semantic decision inside those mechanisms: the phase boundaries, what counts as progress, the watermark that stops a brief posting twice, the choice to keep scheduled work on unattended-safe tools, and the recovery branch itself, which the platform invites you to write and never writes for you.
Keep that branch small and expressed against your own store, the way next_step is, and the portable pattern underneath, reenter a handler and resume from durable progress, is a handful of lines you could carry to any durable-execution runtime. An agent that starts without a caller, survives losing one, and survives losing its own process is worth building. Just be sure you know which of the three you actually turned on.