Your agent is four lines away from a governed cloud endpoint
The lift from a laptop agent to a governed, identity-bearing Foundry endpoint is four lines of hosting code, one YAML file, and five commands, and the contract underneath it is only four requirements long.

The hosted-agent contract on Microsoft Foundry is four requirements long, and a library satisfies all four for you. Here is the whole lift, from an empty directory to a live URL with its own Microsoft Entra identity.
You expect a week of infrastructure work before your agent is reachable from anything but your laptop. It is four lines and one YAML file, and you can finish before your coffee goes cold.
In this article: You will learn exactly what Foundry asks of a container before it will host your agent loop, which package satisfies that contract so you never write an HTTP server, how
azure.yamldecides which endpoints go live, and the fiveazdcommands that take you from empty directory to deployed agent. By the end you will have a version-routed endpoint with its own identity, and a clear accounting of what that cost you.
Most platform quickstarts bury the contract. You follow the steps, something deploys, and you still cannot answer the only question that matters: what does the platform actually require of my code?
For Foundry hosted agents the answer is unusually small: four requirements, and you write none of the code that meets them, because Foundry ships protocol libraries whose whole job is to be the HTTP server so you are not.
What you get back is out of proportion to the lift: a dedicated endpoint, a dedicated Microsoft Entra identity minted per agent at deploy time, per-session isolation, version routing, and log streaming. This article walks the whole crossing with a real example, a supplier-risk desk that reads a chunk of a supplier's public filing against instructions written by a procurement lead. No tools, no memory, no retrieval, no schedule. Just the smallest thing that can honestly be called an agent, running on a governed substrate.
The contract is four requirements long
Take the conclusion before the detail: a hosted agent is an ordinary HTTP container that meets four requirements.
Listen on port 8088. Return 200 OK from GET /readiness. Implement at least one protocol endpoint. Read the environment variables the platform injects [VERIFIED-LEARN: agents-concepts-hosted-agent-contract.md, agents-how-to-init-agent-project.md]. Those four are the whole surface.

Gotcha: the port is 8088, not the 8080 that nearly every other agent-hosting platform defaults to, and the health probe is GET /readiness, not /health or /ping [VERIFIED-LEARN: agents-concepts-hosted-agent-contract.md]. The adapter registers both for you, which is precisely why a hand-rolled server fails in a way the samples never show you. If you are wrapping an existing FastAPI app by hand, those two lines are the ones you will forget.
The package that implements the contract for you
For the Responses protocol, the Python package is azure-ai-agentserver-responses and the .NET package is Azure.AI.AgentServer.Responses [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md, agents-concepts-hosted-agent-contract.md]. The library handles routing, streaming with server-sent events, background execution, cancellation, caching, and the full response lifecycle, plus the /readiness endpoint and the port [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md]. You implement one handler.
Two conveniences inside it do real work. A request context object hydrates the incoming request and the prior conversation turns, so your handler reads plain text rather than parsing a Responses API payload [VERIFIED-LEARN: agents-concepts-hosted-agent-contract.md], and a default_fetch_history_count option decides how many of those turns you get. The platform is managing conversation history on your behalf, which is worth watching, because conversation is not session and the difference costs people data.
One wrinkle in the native path. For Microsoft Agent Framework, the documented entry point is not the protocol library directly but a bridge package, agent-framework-foundry-hosting, described in the migration guide as exactly that [VERIFIED-LEARN: agents-how-to-migrate-hosted-agent-preview.md]. It gives you ResponsesHostServer for /responses and InvocationsHostServer for /invocations [VERIFIED-LEARN: how-to-develop-framework-hosted-agents.md].
CrewAI, Semantic Kernel, and hand-written code use the protocol libraries directly, and LangGraph can too, although the LangChain track also ships its own bridge in langchain_azure_ai.agents.hosting [VERIFIED-LEARN: agents-how-to-migrate-hosted-agent-preview.md, how-to-develop-langchain-hosted-agents.md]. Same contract, same wire, one extra convenience layer on the native track.

The supplier-risk desk, version one
Install the framework and the bridge:
pip install -U agent-framework agent-framework-foundry-hosting azure-identity python-dotenv
Then the whole agent. Everything above ResponsesHostServer is ordinary Agent Framework code; everything from ResponsesHostServer down is the lift.
import os
from agent_framework import Agent
from agent_framework.foundry import FoundryChatClient
from agent_framework_foundry_hosting import ResponsesHostServer
from azure.identity import DefaultAzureCredential
from dotenv import load_dotenv
load_dotenv()
client = FoundryChatClient(
project_endpoint=os.environ["FOUNDRY_PROJECT_ENDPOINT"], # ①
model=os.environ["MICROSOFT_FOUNDRY_MODEL_DEPLOYMENT_NAME"], # ②
credential=DefaultAzureCredential(), # ③
)
agent = Agent(
client=client,
instructions=( # ④
"You are a supplier-risk analyst for a procurement team. "
"Given an excerpt from a supplier's public filing, summarize the "
"financial and operational risks it discloses. Quote the filing for "
"every claim. If the excerpt does not support a conclusion, say so."
),
# The hosting layer manages conversation history, so the model call
# should not also store it.
default_options={"store": False}, # ⑤
)
ResponsesHostServer(agent).run() # ⑥
① The project endpoint is platform-injected into hosted containers, so the code reads it rather than hard-coding a URL.
② The model deployment name is the one variable you declare yourself. It has to match the name you set in azure.yaml.
③ DefaultAzureCredential resolves to your developer login locally and to the agent's own Entra identity once deployed. The code does not change between the two.
④ The instructions are the whole behavior of version one. No tools, no retrieval, no memory, just a policy written in prose.
⑤ Turning off model-level storage keeps conversation history in exactly one place, the hosting layer.
⑥ This single line is the lift. It binds port 8088, exposes POST /responses, and serves the readiness probe.
Note: The full extracted listing at
code/foundry-hyperscaler/part-2-your-first-hosted-loop/listings/02-supplier-risk-desk-main.py
is the same main.py with the marker comments removed.
The last line is the entire lift. The server starts, binds to port 8088, exposes POST /responses, and serves the readiness probe [VERIFIED-LEARN: how-to-develop-framework-hosted-agents.md]. Note store: False: the docs are explicit that the hosting infrastructure manages conversation history, so a model-level store would be a second, redundant copy [VERIFIED-LEARN: how-to-develop-framework-hosted-agents.md].
Note also what the agent cannot do. With no tools, it can only reason over text a caller hands it. An analyst pasting a filing excerpt is a real workflow and a fine place to start, but "go find the filing" is a tool, and tools on this platform become a governed, centrally administered resource rather than a constructor argument. Resist wiring one in here. You would only have to unwire it.
Gotcha: the documentation uses four different names for the model deployment variable across its own pages: AZURE_AI_MODEL_DEPLOYMENT_NAME, MODEL_DEPLOYMENT_NAME, FOUNDRY_MODEL_NAME, and MICROSOFT_FOUNDRY_MODEL_DEPLOYMENT_NAME [CONTESTED: how-to-develop-framework-hosted-agents.md against agents-how-to-author-azure-yaml.md against agents-concepts-hosted-agent-contract.md]. Only one of those matters: the one you declare in azure.yaml and read in your code, because they have to match. This article uses MICROSOFT_FOUNDRY_MODEL_DEPLOYMENT_NAME, which is what azd ai agent init records in your azd environment according to the azure.yaml authoring page [VERIFIED-LEARN: agents-how-to-author-azure-yaml.md]. Do not assume the platform injects a model name for you. FOUNDRY_PROJECT_ENDPOINT it does inject; a model deployment name is not in the platform-injected table [VERIFIED-LEARN: agents-how-to-configure-hosted-agent-env-variables.md].
Set up the toolchain once
Everything after this point runs through the Azure Developer CLI. The Foundry Dev Pack installs azd plus the Foundry extensions in one command: brew install --cask microsoft/foundry/devpack && foundry-devpack install on macOS, winget install Microsoft.FoundryDevPack on Windows, and a curl script from aka.ms/foundry-devpack-install.sh on Linux [VERIFIED-LEARN: includes-how-to-develop-install-cli-sdk-content-foundry.md]. If you already run azd, take the extension bundle directly:
azd ext install microsoft.foundry
azd auth login
microsoft.foundry is a meta-package pulling in every Foundry extension: azure.ai.agents for agents, plus separate extensions for connections, the inspector, projects, routines, skills, and toolboxes, each independently versioned so you can upgrade one at a time [VERIFIED-LEARN: agents-how-to-install-cli-foundry-extensions.md].
Two permissions matter before you go further. Deploying a hosted agent needs Foundry Project Manager at project scope; creating a brand-new Foundry project needs Owner at resource-group scope [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md, agents-quickstarts-quickstart-hosted-agent.md]. Merely calling the finished agent needs only Foundry Agent Consumer [VERIFIED-LEARN: agents-includes-environment-setup-content.md], which is the role your analysts eventually get and nothing more.
Gotcha: the hosted-agent runtime is generally available, but the azd ai tooling around it is not. The extension installation page, the project-init page, the azure.yaml authoring page, and the protocol-adapter page all carry the preview banner [VERIFIED-LEARN: agents-how-to-install-cli-foundry-extensions.md, agents-how-to-init-agent-project.md, agents-how-to-author-azure-yaml.md, agents-how-to-add-protocol-adapter.md] [PERISHABLE: as of September 2026]. The version floors move too: the extensions page asks for azd 1.25.2 or later while the azure.yaml page and the quickstart both ask for 1.27.1 [CONTESTED]. Take the higher number and upgrade when a command behaves strangely.
azure.yaml is where the endpoint is declared
A hosted agent project keeps all of its configuration in one azure.yaml at the project root [VERIFIED-LEARN: agents-how-to-author-azure-yaml.md]. This is the file worth understanding, because it is where you say which protocol endpoints go live.
Each entry under services has a host field naming the kind of Foundry resource it declares, and services reference each other through uses, the dependency graph azd resolves at provision and deploy time [VERIFIED-LEARN: agents-how-to-author-azure-yaml.md]. A minimal project has two: a project that owns the model deployment, and an agent that depends on it.
# yaml-language-server: $schema=https://raw.githubusercontent.com/Azure/azure-dev/main/schemas/v1.0/azure.yaml.json
name: supplier-risk-desk
services:
ai-project:
host: azure.ai.project # ①
deployments:
- name: gpt-5.4-mini # ②
model:
format: OpenAI
name: gpt-5.4-mini
version: "2026-03-17"
sku:
name: GlobalStandard
capacity: 10
supplier-risk-desk:
host: azure.ai.agent # ③
kind: hosted
project: src/supplier-risk-desk
language: docker
uses: # ④
- ai-project
name: supplier-risk-desk
description: Summarizes risk disclosures in supplier filings.
protocols: # ⑤
- protocol: responses
version: 2.0.0
env:
MICROSOFT_FOUNDRY_MODEL_DEPLOYMENT_NAME: ${MICROSOFT_FOUNDRY_MODEL_DEPLOYMENT_NAME} # ⑥
startupCommand: python main.py # ⑦
container:
resources:
cpu: "0.25"
memory: 0.5Gi
① The host value names the kind of Foundry resource this service declares. A project resource is what owns model deployments.
② The deployment name declared here is the name your code has to read at runtime, which is why the two sides must agree.
③ The agent service declares a hosted agent, the container Foundry runs for you.
④ uses is the dependency edge. It tells azd to provision the project before the agent that needs its model.
⑤ The protocols list decides which endpoint URLs the platform publishes. Declare only what your handler actually implements.
⑥ The ${VAR} placeholder resolves from the active azd environment, so one file serves dev, staging, and production.
⑦ Both the local run and the deployed container read startupCommand, so your agent starts exactly one way in either place.
Note: The full extracted listing at code/foundry-hyperscaler/part-2-your-first-hosted-loop/listings/03-azure-yaml.yaml is the same manifest with the marker comments removed.

The protocols block is the load-bearing part. The values you put there determine which endpoint URLs the platform publishes. Valid protocols are responses, invocations, and a2a, with invocations_ws for WebSocket and activity for Teams and Microsoft 365 [VERIFIED-LEARN: agents-concepts-azure-yaml-reference.md]. One agent can declare several. Declaring a protocol you have not implemented gets you a live URL that fails, so the list here and the handlers in your code are one decision made twice.
Three more fields deserve a sentence each. container.resources accepts cpu from "0.25" to "4.0" and memory from 0.5Gi to 8.0Gi [VERIFIED-LEARN: agents-how-to-author-azure-yaml.md], which is the honest ceiling on a single hosted agent's compute. startupCommand is used both by the local run and by the deployed container [VERIFIED-LEARN: agents-how-to-run-hosted-agent-locally.md], so there is exactly one way your agent starts. And ${VAR} placeholders resolve from the active azd environment in .azure/<env>/.env at run and deploy time [VERIFIED-LEARN: agents-concepts-cli-agent-development.md].
Gotcha: do not declare FOUNDRY_PROJECT_ENDPOINT in the env map. The platform injects it into hosted containers automatically, and azd ai agent run sets it locally from your azd environment, so declaring it is redundant and risks shadowing the platform-managed value [VERIFIED-LEARN: agents-how-to-configure-hosted-agent-env-variables.md]. The entire FOUNDRY_* and AGENT_* prefix space is reserved [VERIFIED-LEARN: agents-how-to-configure-hosted-agent-env-variables.md]. Name your own variables something else.
In production: secrets do not belong in the env map as literals, and they do not belong in the Dockerfile either, where they are fixed at build time and visible to anyone with the image. Store them in a Foundry project connection and reference them with a ${{connections.<name>.credentials.<field>}} placeholder, which the platform resolves at sandbox start [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md]. A GET on the agent version returns the literal placeholder text, never the resolved secret.
The azd ritual, five commands
This is the sequence worth running once with your hands rather than reading.

Scaffold. Run this in the directory that already holds the main.py above:
azd ai agent init --deploy-mode code
The CLI detects your files and generates an azure.yaml service entry around them without overwriting anything [VERIFIED-LEARN: agents-how-to-init-agent-project.md]. The wizard prompts for the agent name, the project, the tenant, the subscription, the region, the model, its version, its SKU, the capacity, and the deployment name [VERIFIED-LEARN: agents-quickstarts-quickstart-hosted-agent.md]. Point -m at a sample's azure.yaml instead to adopt that manifest and download its source.
Provision. Create the Azure side of the project:
azd provision
This creates a resource group holding, among other things, a Foundry instance, a project with a model deployment, an Application Insights instance, and a container registry for agent images [VERIFIED-LEARN: how-to-develop-framework-hosted-agents.md].
Run locally, against the real project. This is the step that makes the inner loop tolerable:
azd ai agent run
The command auto-detects the project type, installs dependencies, starts your agent on localhost:8088 using the startupCommand from azure.yaml, and opens the agent inspector in your browser. It also injects your azd environment, so FOUNDRY_PROJECT_ENDPOINT and every value you set with azd env set are present without manual configuration [VERIFIED-LEARN: agents-how-to-run-hosted-agent-locally.md, agents-quickstarts-quickstart-hosted-agent.md]. Your container is on your laptop, and the model, the project, and your Entra credential are the real ones. Pass --port when 8088 is taken.
Invoke it. From a second terminal:
azd ai agent invoke --local "Summarize the supply-chain risks in this excerpt: ..."
Or bypass the CLI entirely and speak the protocol, which is what the caller on the other side of the endpoint will eventually do:
curl -X POST http://localhost:8088/responses \
-H "Content-Type: application/json" \
-d '{"input": "Summarize the supply-chain risks in this excerpt: ...", "stream": false}'
Set stream to true and the same handler emits Responses API server-sent events such as response.created, response.output_text.delta, and response.completed [VERIFIED-LEARN: how-to-develop-framework-hosted-agents.md]. You wrote none of that.
Deploy.
azd deploy
During deploy, the CLI builds your container image remotely in Azure Container Registry so you do not need local Docker, pushes it, creates a hosted agent version on Foundry Agent Service, and creates a dedicated Microsoft Entra agent identity with the RBAC roles the agent needs to reach models and tools [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md]. Each azd deploy creates a new version, previous versions are preserved, and the latest is active by default [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md]. Use azd up when you are changing infrastructure and code at once; it is azd provision plus azd deploy [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md].
The Entra identity is not a footnote. The platform mints a service principal per hosted agent at deploy time, and your container authenticates as itself rather than as a shared credential [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md]. It is the single most consequential thing you got for free here.
Gotcha: whether you need Docker running depends on a choice the docs describe inconsistently. Code deployment, the default for Python and .NET, uploads your source as a ZIP and builds it remotely, and the azure.yaml authoring page states plainly that it does not require Docker [VERIFIED-LEARN: agents-how-to-author-azure-yaml.md, agents-how-to-init-agent-project.md]. The Agent Framework hosting page says Docker must be running locally because azd ai agent run builds the image from the sample's Dockerfile [CONTESTED: how-to-develop-framework-hosted-agents.md against agents-how-to-author-azure-yaml.md]. Both are true of different deploy modes. If you pass --deploy-mode code, expect a virtual environment; if you use container mode with a local build, start Docker. One more trap if you do build locally: the platform requires x86_64 images, so on Apple Silicon build with docker build --platform linux/amd64 . [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md].
Calling it like any OpenAI endpoint
The deploy output prints two links, one to the agent playground in the portal and one to the agent endpoint [VERIFIED-LEARN: agents-quickstarts-quickstart-hosted-agent.md]. The endpoint is live the moment the agent exists, with no separate publish step, and its URL does not change as you roll out new versions. Version routing is a property of the agent, not of the URL: all traffic goes to the newest version by default, and you can pin it to a specific one in production [VERIFIED-LEARN: agents-how-to-configure-agent.md].
Each declared protocol gets its own path under that endpoint. For Responses:
https://{account}.services.ai.azure.com/api/projects/{project}/agents/{agent}/endpoint/protocols/openai/responses
The agent configuration reference verifies that shape alongside the Activity, Invocations, A2A, and MCP paths [VERIFIED-LEARN: agents-how-to-configure-agent.md]. Because the protocol is OpenAI-compatible by default [VERIFIED-LEARN: agents-how-to-agent-applications.md], that URL is a base_url for an ordinary OpenAI client holding an Entra bearer token. The caller needs only Foundry Agent Consumer on the project or agent scope, and API key authentication is not supported on agent endpoints at all [VERIFIED-LEARN: agents-how-to-configure-agent.md].
Two operational commands close the loop. azd ai agent show prints the name, version, protocols, container resources, environment variables, and creation timestamp [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md]. azd ai agent monitor --follow streams container logs while you interact with the agent [VERIFIED-LEARN: agents-quickstarts-quickstart-hosted-agent.md], which is the first thing to reach for when a deployed agent behaves differently from the local one.
One quiet thing happened during azd provision. The platform injects APPLICATIONINSIGHTS_CONNECTION_STRING into your container, and when an Application Insights resource is attached to the project, OpenTelemetry tracing is on by default and traces may contain personal data and customer content [VERIFIED-LEARN: agents-how-to-deploy-hosted-agent.md]. Provision through azd and you have that instance whether you asked for one or not. Wire the project by hand and you probably do not, and your traces dashboard will be silently empty.
Do this today
- Install the toolchain:
azd ext install microsoft.foundryfollowed byazd auth login, and confirm yourazdversion is at least 1.27.1. - Check your role assignments before you write code. Deploying needs Foundry Project Manager at project scope, and creating a fresh project needs Owner at resource-group scope.
- Write the four lines of hosting code around whatever agent you already have, then run
azd ai agent initinside that directory and read theazure.yamlit generates. - Run
azd provisionand thenazd ai agent run, and invoke the agent onlocalhost:8088against the real project, the real model, and your real credential. - Deploy with
azd deploy, then runazd ai agent showand read the version, the protocols, and the Entra identity the platform minted for you.
What just transferred, and what it cost
Foundry took the HTTP server, the health probe, the TLS termination, the container build, the image registry, the rollout, the version routing, the agent's Entra identity, and the conversation history the Responses protocol hydrates into your handler.

What stayed yours is everything in main.py above the last line: the instructions, the model choice, the stopping conditions, and the judgment about what the agent should refuse to conclude from a filing excerpt.
The honest accounting of what this cost you is small, which will not remain true. Your code imports one Foundry-specific package and reads one platform-injected environment variable. Swap ResponsesHostServer for a plain HTTP server and this agent runs anywhere. The lock-in worth tracking starts later, when conversation history, memory, and toolbox configuration live in Foundry and have no bulk export. Enjoy the cheap seat while you are in it.
Run the five commands. Paste in a real filing excerpt, read what comes back, and notice that you have a governed, identity-bearing, version-routed endpoint and you never wrote a line of infrastructure. Then notice the second thing, the one that should keep you up: the answer quality is entirely down to those instructions, the only part of the system nobody can sell you.