Microsoft Foundry Is Not a Product You Log Into. It Is an ARM Resource.
Skip the feature tour. One structural fact, that a Microsoft Foundry resource is an ordinary ARM resource under a shared provider namespace, explains most of how the platform governs, authenticates, and scales your agents, and it draws the line between the loop you keep and the substrate you rent.

Most guides open with a feature tour. Start instead with the one structural fact that makes the rest of the platform legible: your existing Azure governance already applies, because this is the same kind of resource your security team has been governing all along.
You already know how to write an agent loop. The question nobody answers honestly is what the platform actually owns, and what adopting it costs you later.
In this article: You will learn the single fact that explains most of Microsoft Foundry's behavior, why the resource and project split decides your isolation design, how one endpoint and one credential replaced a pile of per-service URLs, the difference between a prompt agent and a hosted agent, and the responsibility line between the agentic loop you keep and the substrate you rent. You will also learn the provenance labels this series uses to date its facts, because Microsoft's own documentation disagrees with itself.
You have already written an agent loop. It runs on your laptop, or in a container you operate, and it works. The question in front of you now is a different one. Your organization has standardized on Azure, and somebody has to decide what the platform owns and what your team keeps owning at three in the morning.
Most guides answer that with a feature tour. This one starts somewhere less glamorous and considerably more useful.
A Microsoft Foundry resource is an Azure Resource Manager resource of type Microsoft.CognitiveServices/accounts with kind AIServices, and a Foundry project is a subresource of it, Microsoft.CognitiveServices/accounts/projects [VERIFIED-LEARN: includes-resource-provider-kinds.md]. That is not trivia. It is the reason your existing Azure Policy definitions keep working, the reason governance here is Azure governance rather than a product-specific console, and the reason an Azure OpenAI resource can be upgraded in place without losing its endpoint.
The through-line of this series is simple to state and expensive to forget: the loop is yours, the substrate is theirs. You keep your agentic loop in whatever framework you chose, and Foundry supplies the managed harness around it.
How to read the labels in this series
Microsoft's AI platform renamed itself twice in two years, ships preview features next to generally available ones on the same page, and carries pages that disagree with each other about catalog sizes and token scopes. A guide to a platform moving this fast either dates its facts or misleads its readers. This one dates them inline.

[VERIFIED-LEARN]means the official documentation mirror this series is written against confirms the claim, and the label names the file.[VERIFIED-BOTH]means the mirror confirms it and independent research corroborates it.[RESEARCH-ONLY]means it is attested elsewhere but absent from the mirror. Check a live Microsoft Learn page before you act on it.[CONTESTED]means sources disagree, sometimes two Microsoft pages with each other. The disagreement gets stated rather than silently resolved.[PERISHABLE]means it is true on the stated date and expected to change.[REFUTED]means somebody claims it and the documentation shows it is wrong.
The rule behind all six: prefer the Learn feature table over a conference announcement when they disagree. An announcement describes intent. A preview banner describes what you can rely on next quarter.
Labels go on facts you would act on, such as preview status, quotas, region lists, retirement dates, and API names. They do not go on opinions. Everything here is Python, because the bring-your-own hosted-agent path is documented for Python and C# only, and the LangChain track is Python-only [VERIFIED-BOTH].
The resource model is the load-bearing fact
Foundry organizes AI workloads in three layers: a top-level resource for governance, projects underneath it for development isolation, and connected Azure services for storage, search, and secrets [VERIFIED-LEARN: concepts-architecture.md].

The resource is where networking, encryption, policy, and role assignments live. The project is where agents, evaluations, and files live. One resource carries many projects, which is how a platform team gives each product team an isolated workspace on shared, governed infrastructure.
The interesting part is the namespace. Foundry shares Microsoft.CognitiveServices with Azure OpenAI, Speech, Vision, and Language [VERIFIED-LEARN: includes-concepts-architecture-1.md]. Resource types under the same provider namespace share management APIs and use similar Azure RBAC actions, networking configuration, and Azure Policy aliases. The documentation draws the conclusion directly: if you upgrade from Azure OpenAI to Foundry, your existing custom Azure policies and RBAC actions continue to apply [VERIFIED-LEARN: concepts-architecture.md], and the upgrade preserves your endpoint, API keys, and existing state [VERIFIED-LEARN: what-is-foundry.md].
Read that as an adoption argument, not a feature. Private Link, customer-managed keys, Azure Policy, RBAC, diagnostic settings, and resource locks all behave the way they behave everywhere else in your subscription. Your security team does not need to learn a new governance console. Compare that to an agent platform living entirely outside your cloud's control plane, and the difference is an entire compliance review.
Two terms, defined once. The control plane is resource management: creating resources and projects, assigning roles, rotating keys, configuring Private Link. The data plane is runtime: chat completions, agent invocations, evaluation jobs. Azure authorizes the first with RBAC actions and the second with RBAC dataActions, and Foundry keeps the distinction sharp [VERIFIED-LEARN: includes-concepts-authentication-authorization-foundry-content.md]. When this series says "control plane," it means that, not a dashboard.
One endpoint, one credential
Every SDK here targets the same address: the project endpoint, shaped https://<project>.services.ai.azure.com [VERIFIED-LEARN: how-to-navigate-from-classic.md]. That one endpoint replaced the old pile of per-service URLs for openai, azureml, cognitiveservices, search, and speech. Stale multi-endpoint URLs are the documented cause of connection failures during migration.
Authentication defaults to Microsoft Entra ID. The shape below is the one you will write a hundred times: a standard OpenAI client pointed at the project's OpenAI-compatible route, holding a bearer token that DefaultAzureCredential resolves from your environment, your managed identity, or your Azure CLI login.
from openai import OpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
client = OpenAI(
base_url="https://my-project.services.ai.azure.com/openai/v1",
default_headers={"Authorization": f"Bearer {get_bearer_token_provider(DefaultAzureCredential(), 'https://cognitiveservices.azure.com/.default')()}"}
)
Nothing there is Azure-specific except the base URL and the credential, which is the point. The Azure-specific client class is gone, and swapping base_url is most of a migration in either direction. Hold onto that, because it is the foundation of the lock-in argument below.

Gotcha: the documentation gives two different Entra token scopes. The authentication concept page says bearer tokens are scoped to https://ai.azure.com/.default, and its troubleshooting table tells you to verify exactly that scope on a 401, while the migration sample above passes https://cognitiveservices.azure.com/.default [CONTESTED: includes-concepts-authentication-authorization-foundry-content.md against how-to-navigate-from-classic.md]. Both pages are official. Default to https://ai.azure.com/.default, which 51 pages in the mirror use against 10 for the other, and which is what the page whose whole job is authentication documents and troubleshoots. If you hit a 401 that your role assignment does not explain, try the other scope before you rebuild anything, and prefer DefaultAzureCredential over hand-built headers so the library owns that detail.
API keys still exist, and for agents they are not an option. The feature support matrix marks Agents service, Evaluations, and Toolbox as Entra-only, with API keys explicitly unsupported [VERIFIED-LEARN: includes-concepts-authentication-authorization-foundry-content.md]. A key cannot express user identity, which is precisely what the agent surface needs.
Authorization uses built-in roles: Foundry Agent Consumer, Foundry User, Foundry Project Manager, Foundry Account Owner, and Foundry Owner [VERIFIED-LEARN: concepts-rbac-foundry.md]. Agent Consumer can invoke agent endpoints and nothing else, Foundry User can build in a project but cannot create one, and only the two owner roles cross both planes. That granularity is the enterprise AI agent governance story Foundry leads on, and it is ordinary Azure RBAC rather than a bespoke permissions model.
Two agent types, and everything downstream depends on the difference
Foundry has exactly two agent types [VERIFIED-BOTH].
A prompt agent is declarative. You supply instructions, a model, and tools through the portal or the SDK, and Foundry runs it. No application code, no container, no runtime to maintain [VERIFIED-LEARN: agents-overview.md].
A hosted agent is your code. You build it with Microsoft Agent Framework, LangGraph, Semantic Kernel, the OpenAI Agents SDK, the Anthropic Agent SDK, the GitHub Copilot SDK, or nothing at all, and you ship it as a container image or a zip of source that Foundry builds for you. The platform runs it with a managed endpoint, automatic scaling, a dedicated Microsoft Entra identity per agent, session-level state persistence, and end-to-end observability [VERIFIED-LEARN: agents-overview.md]. Foundry hosted agents are generally available [VERIFIED-LEARN: how-to-navigate-from-classic.md].

Hosted agents are where the seam between your loop and their substrate is visible. The prompt agent is genuinely useful and genuinely limited: the right answer when instructions and tools are enough, the wrong answer the moment you need custom orchestration or your own dependencies.
One boundary the docs draw that the industry usually blurs. Hosted agents run in a Microsoft-managed, VM-isolated sandbox, one per session, and workloads run in logically isolated environments per Foundry resource with no shared runtime containers between tenants [VERIFIED-LEARN: concepts-architecture.md, agents-overview.md]. The documentation deliberately does not name the underlying compute product, and neither will this series. Vendor research confidently asserts Azure Container Apps, another source denies it, and the hosted-agent pages mention no such thing [CONTESTED]. A subnet delegation to Microsoft.App/environments does appear in the bring-your-own-network requirements [VERIFIED-LEARN: concepts-architecture.md], which is networking plumbing, not a product name. Design against the contract, not against a guess about what is underneath it.
Three ways to get a harness
The decision you are actually making is older than this platform.
Build it. You write the loop, the context manager, the validators, the memory tiers, the sandbox, and the traces. Maximum control, maximum surface area, and every component becomes something your team operates.
Rent the loop. You hand the entire agentic loop to a vendor, define an agent and an environment, and watch a server-side loop through an event stream. You configure a harness instead of coding one.
Rent the substrate. You keep the loop in your own framework, in your own repo, under your own tests, and rent everything around it: hosting, identity, isolation, tool governance, state, memory, tracing, and evaluation.
This series is the third path, on Azure. All three hyperscalers converged on it at roughly the same time, which tells you the market read the same signal. The harness is infrastructure, infrastructure is a business, and the loop belongs to the customer. Amazon sells Bedrock AgentCore, Google sells Vertex AI Agent Engine, Microsoft sells Foundry. Most of what you learn on one transfers, because the responsibility split transfers even when the SKUs do not.
Foundry earns deep treatment for two reasons. Its bring-your-own-framework path is explicitly first-class and framework-agnostic, with the protocol libraries documented as working with Agent Framework, LangGraph, Semantic Kernel, or hand-written code [VERIFIED-BOTH]. It also carries an enterprise identity and governance story further up the stack than its peers. Note the status honestly: Foundry Control Plane is preview, not generally available [VERIFIED-LEARN: control-plane-overview.md, how-to-navigate-from-classic.md], despite at least one research source claiming a March 2026 GA. That is the "prefer the feature table" rule doing its first piece of real work.
Two loops, one substrate
Every platform topic in this series wires a Foundry service into two different agent frameworks, in the same order, every time.
Microsoft Agent Framework on Azure is the native path, Microsoft's stated successor to Semantic Kernel and AutoGen, and it internalizes the harness. Sessions, middleware, tool approval, and memory are first-class APIs inside the framework, and Foundry integrations show up as first-party one-liners.
LangGraph on Azure, with DeepAgents, is the bring-your-own path, and it externalizes every harness concern as a constructor argument you can see and swap. It is not a second-class citizen: the official docs carry a first-class eight-page langchain-azure-ai track covering agents, hosted agents, models, memory, middleware, toolbox, and tracing, and two parallel hosting pages exist, one per framework [VERIFIED-LEARN].
Neither is better. They distribute the same responsibilities differently, and watching one managed service get consumed both ways is what teaches you where the substrate ends and your loop begins. Frameworks churn; that boundary does not.
Three framings that run through everything
Conversation is not session is not memory. Foundry exposes four state services with four different lifetimes: a conversation (a durable message record living independently of compute), a session (a compute lease with a persistent filesystem that ages out on an idle timer), a state store (a durable key-value partition you control), and memory (LLM-extracted long-term knowledge, still in preview [VERIFIED-BOTH] [PERISHABLE: as of September 2026]). Each has its own failure mode, and merging any two of them in your head produces the wrong isolation design and loses data in production.
Lock-in concentrates in the managed state services, not the models. The model catalog is the least sticky thing in Foundry. It is multi-vendor, it carries Anthropic and Meta models alongside the GPT family, and the overview page claims more than 10,000 models [PERISHABLE: what-is-foundry.md, revised August 13, 2026], a number that drifts by the month and that Microsoft's own pages do not agree on. Swapping a model is usually swapping a deployment name. The sticky parts are the managed conversation store, memory, and toolbox configuration, where there is no bulk export and where the Assistants-to-Agents migration tool moves code constructs and explicitly not state data.
Documentation drift is the standing risk. You already saw two official pages disagree about a token scope. The docs also carry two different Model Router model counts in what's-new entries from different dates, which is why this series describes that mechanism and never prints the number [PERISHABLE]. The labels are the mitigation.
The rename chain, handled once
Microsoft's AI platform went Azure AI Studio, then Azure AI Foundry, then Microsoft Foundry, and the services portfolio went Azure Cognitive Services, then Azure AI Services, then Foundry Tools. Through all of it, the Azure resource type stayed Microsoft.CognitiveServices/accounts [VERIFIED-LEARN: how-to-navigate-from-classic.md]. Research places the final rename on November 18, 2025 [RESEARCH-ONLY], so verify that date against a live page before you quote it.
Gotcha: four names will strand you, and the first one will cost you an afternoon.
- Hub-based projects are not visible in the current Foundry portal at all. They live in the Foundry (classic) portal, and new SDK samples do not work against them
[VERIFIED-LEARN: how-to-navigate-from-classic.md]. If your projects vanished, you are in the wrong portal, not the wrong subscription. - Studio and AI Foundry pages still rank in search and describe a different resource model.
- Assistants is retired. The Assistants API sunset on August 26, 2026, replaced by the Responses API, and
azure-ai-inferenceretired the same day in favor of the standardopenaipackage[VERIFIED-BOTH]. azure-ai-projects2.x targets the current portal while 1.x targets the classic one. Mixing a 2.x sample with a 1.x setup produces errors that look like bugs.
One more trap sits next to those. The Responses API is not available in every Azure region, so a Foundry resource in an unsupported region cannot create or run agents at all [VERIFIED-LEARN: how-to-navigate-from-classic.md]. Check the region before you provision, not after.
The running example: a supplier-risk desk
One example grows across the whole series: a supplier-risk desk for a procurement team. It watches supplier news and public filings, cross-checks findings against internal contracts, remembers each buyer's risk thresholds, runs unattended every weekday morning, survives a crash mid-review, fans research out to per-supplier worker agents, and produces a weekly risk brief that gets evaluated before it ships to Teams.
Two properties make it the right vehicle rather than an arbitrary demo. It is multi-tenant by construction, because many buyers share one deployed agent, and multi-tenancy is what exercises Foundry's identity, isolation, and session-multiplexing surface, the part of the platform where a default value can leak one customer's data to another. Its work is also long-running and asynchronous, which exercises the state and resilience surface, where confusing a session with a conversation costs you a day of analysis.
It is also deliberately not a coding agent. Coding agents make every platform look good. A procurement desk reading untrusted supplier filings makes the platform's guardrails, tool governance, and threat model do actual work.
What Foundry owns, and what stays yours
Here is the split the rest of the series enforces, service by service.
| Responsibility | What Foundry owns | What stays yours |
|---|---|---|
| Endpoint and scaling | Managed endpoint per agent, automatic scaling of container instances per session and request volume | The loop inside the container, its stopping conditions, and its cost ceiling |
| Identity | A dedicated Microsoft Entra identity per hosted agent, plus Azure RBAC and the control-plane and data-plane separation | Which roles that identity actually needs, and the entitlement model for your end users |
| State | Conversation durability, session filesystem persistence, and a durable state store as raw material | Checkpoint semantics, retention policy, and the isolation key |
| Tools | A governed registry, credential injection, versioning, and one MCP-compatible endpoint | Tool descriptions, exclusion clauses, structured errors, and which tools an agent should have at all |
| Memory | Extraction, consolidation, and conflict resolution | Saliency, write-time validation, and what must never be remembered |
| Safety | Guardrails at fixed intervention points, enforced outside your code | The domain rules, the escalation policy, and the threat model |
| Quality | Judges, sampling, dashboards, and a CI-runnable evaluation surface | The golden set, judge calibration, which dimensions matter, and the promotion threshold |
The table covers what the platform manages for each responsibility and what remains your engineering work.
- Endpoint and scaling: Foundry owns the managed per-agent endpoint and automatic scaling of container instances per session and request volume. The loop inside the container, its stopping conditions, and its cost ceiling stay yours.
- Identity: Foundry owns a dedicated Microsoft Entra identity per hosted agent, Azure RBAC, and the control-plane and data-plane separation. Deciding which roles that identity needs and how end users are entitled stays yours.
- State: Foundry owns conversation durability and session filesystem persistence, and supplies a durable state store as raw material. Checkpoint semantics, retention policy, and the isolation key stay yours.
- Tools: Foundry owns a governed registry, credential injection, versioning, and a single MCP-compatible endpoint. Tool descriptions, exclusion clauses, structured error responses, and the decision about which tools an agent should have stay yours.
- Memory: Foundry owns extraction, consolidation, and conflict resolution. Saliency, write-time validation, and what should never be remembered stay yours.
- Safety: Foundry owns guardrails enforced at fixed intervention points, outside your code. The domain rules, the escalation policy, and the threat model stay yours.
- Quality: Foundry owns the judges, the sampling, the dashboards, and a CI-runnable evaluation surface. The golden set, judge calibration, which dimensions matter, and the promotion threshold stay yours.

Read the right column again. It is not a list of leftovers. It is the list of controls whose absence causes the failures you have already lived through. Microsoft will sell you an identity per agent and a dashboard of judges. It cannot sell you the entitlement model, or the verdict, or the decision about what your agent should never be allowed to do. Those were always the engineering. The platform removed the scaffolding that was hiding them.
Do this today
Five checks you can run before you write a line of agent code.
- Run
az resource showagainst an existing Azure OpenAI resource and read itstypeandkind. SeeingMicrosoft.CognitiveServices/accountsis the whole adoption argument in one command. - Ask your security team which Azure Policy definitions already target
Microsoft.CognitiveServices. Whatever comes back is governance you do not have to rebuild. - Confirm the Responses API is available in the region you were about to provision in. A Foundry resource in an unsupported region cannot create or run agents at all, and that is a rebuild, not a config change.
- Check which portal you are in before you file a "my projects disappeared" ticket. Hub-based projects live in Foundry (classic) and do not appear in the current portal.
- Pin
azure-ai-projectsto 2.x in a fresh environment and delete everyazure-ai-inferenceimport. The 1.x and 2.x split produces errors that look like bugs, andazure-ai-inferenceis retired in favor of the standardopenaipackage.
Take the infrastructure, spend the time on the right column
There is no medal for hand-rolling a VM-isolated sandbox or a credential-injecting tool gateway. The ones Microsoft operates are better than the ones you were going to build, and they arrive already wearing your subscription's policy, encryption, and network controls, because they are ordinary Azure resources.
Take them. Then spend the time you just saved on the right column of that table, which is the only column your users will ever notice.
The fact to carry forward is the small one from the top. Foundry is not a console you log into. It is a resource type in your subscription, governed by rules you already wrote, running a loop you still own. Everything else is a consequence of that sentence.
This is Part 1 of "Microsoft Foundry as a Hyperscaler Control Plane," a 17-part guide to running your own agent loop on Azure's managed agent substrate.