Microsoft Foundry is a platform for building and operating agentic AI applications. Foundry starts with the widest model selection on any cloud — models from Microsoft, OpenAI, Anthropic, Meta, Mistral, DeepSeek, Hugging Face, and others, spanning frontier, open-source, and custom weights — all accessible through a single endpoint and a single set of SDKs in Python, C#, JavaScript, and Java.
On top of those models sits the Foundry Agent Service: multi-agent orchestration with built-in memory, knowledge grounding through Foundry IQ, and a catalog of connectable tools via agentic protocols, so agents can work with enterprise data. Once agents are running, Foundry provides end-to-end tracing, real-time monitoring, continuous evaluations, and a prompt optimizer that improves agent behavior based on eval results — observability and quality loops that are part of the platform.
Alongside pay-per-token (lowest-friction path to get started) and provisioned throughput (predictable, high-performance production workloads on frontier models), Foundry Managed Compute is the third deployment option in Foundry: a managed GPU platform-as-a-service for open-source and custom models.
You deploy a model instance described by the things that matter to your workload — parameter count, context length, and whether you want to optimize for latency or throughput — and Foundry handles the GPU topology underneath, whether the instance lands on one accelerator or several, so you think and plan in model terms.
cpp — without redeploying your model, while model configuration, deployment behavior, and routing stay with you.
Source link







