Keeping AI inference on-premises does not settle every sovereignty question. Teams still need to decide who can administer the system, which dependencies require external connectivity, and how much work it would take to replace a model or move a workload.
Microsoft’s October 5, 2026 announcement introduces a Sovereign AI white paper developed with contributions from NVIDIA. The paper covers architecture and operating decisions; it does not announce a new AI model or a newly launched cloud service.
The framework organizes those decisions around four principles: control, choice, flexibility and resilience. For regulated organizations and public-sector teams, its value lies in translating those principles into requirements that engineers can implement and test. Microsoft’s accompanying Foundry Local documentation provides a concrete deployment path for some of those requirements, although the Azure Local offering remains in preview.
Start With the Workload, Not the Deployment Label
Microsoft defines sovereign AI as designing, deploying and operating AI workloads under defined controls for data, access, governance, infrastructure and operations. In this framing, sovereignty is a workload-specific risk-management question, not a single infrastructure configuration.
A system handling public documents may have different requirements from one processing sensitive records or supporting an operational decision. Buying the same platform for both does not establish that their controls are appropriate.
Microsoft recommends starting with four practical questions:
- What data will the workload use or generate, and where may processing occur?
- Who needs access, including administrative access?
- What must continue working when connectivity fails?
- How easily must models or infrastructure be replaceable?
The company says public cloud environments can meet the requirements of many workloads. Others require greater operational control, infrastructure under the organization’s authority, or limited-connectivity operation.
Documenting those requirements should come before selecting the deployment environment. “Sovereign” should describe the controls a workload actually needs, not substitute for that analysis.
Four Principles Become Four Architecture Decisions
Microsoft’s framework is broad. The following questions translate it into decisions a platform team can investigate. They are an architectural reading of the announcement, not a separate certification checklist.
| Principle | Concrete architecture question |
|---|---|
| Control | Where may prompts, retrieved data and outputs be processed, and who may administer each component? |
| Choice | Which models, runtimes and infrastructure options must remain available to the organization? |
| Flexibility | What can change without redesigning the application or its governance arrangements? |
| Resilience | Which functions must survive a connectivity disruption, and what dependencies could prevent that? |
Control extends beyond stored data. Microsoft explicitly includes prompts, models, agents, outputs and the systems operating them in its sovereignty discussion. A database’s location cannot answer every question about an AI application. Teams should map the request path, including retrieval, inference and operational logging, and identify the administrative authority over each step.
Choice concerns both models and infrastructure. Microsoft argues for separating the AI platform from a particular model so organizations can evaluate alternatives as capabilities and requirements change. NVIDIA’s contribution spans accelerated computing, AI software and model options. These are vendor-described capabilities, not evidence that every combination is interchangeable or suitable for every regulated workload.
Flexibility requires examining the cost of change. An organization might need to move processing closer to its data, replace a model, or adopt different hardware. A shared API can reduce application changes, but migration still requires checking model behavior, capacity, authentication and operational procedures.
Resilience needs a defined service boundary. Keeping a model server running is only one part of keeping an application available. If retrieval, identity or another required service remains external, the application may still fail when the connection disappears.
Microsoft recommends identifying which functions must continue, which can pause, and what services they depend on. The stronger interpretation is to turn that list into failure tests. “Local inference” alone is no proof of continuity.
Foundry Local Provides an On-Premises Inference Option
Microsoft’s Foundry Local on Azure Local overview describes enterprise-scale inference on an Azure Arc-enabled Kubernetes cluster. Azure Local is the validated and supported platform for this deployment model.
The deployment targets organizational infrastructure, Kubernetes operations and centralized model serving. It is distinct from the device-oriented Foundry Local option for embedding AI in applications running on end-user hardware.
Documented capabilities include OpenAI-compatible generative request patterns, CPU- and GPU-backed execution, multi-model serving and disconnected operation. Microsoft also documents predictive model serving, so the platform is not limited to chat-style applications.
The architecture separates model definitions from active deployments. A Kubernetes inference operator manages lifecycle changes through declarative Model and ModelDeployment resources. Models can come from the Foundry catalog or an organization’s own registry; deployment resources describe runtime requirements such as scaling and endpoint exposure.
For generative inference, Microsoft lists ONNX-GenAI for CPU or GPU execution and vLLM for GPU-only, high-throughput scenarios. Hardware selection needs to follow model size, concurrency and latency requirements. Teams should not assume that every workload requires the same NVIDIA configuration.
OpenAI-compatible requests can help application integration, but they do not establish identical output quality or full behavioral compatibility between models. Switching models still needs workload-specific evaluation.
Microsoft labels Foundry Local as preview and warns that features, approaches and processes can change or have limited capabilities before general availability. The white paper does not change that status.
Connected and Disconnected Deployments Have Different Dependencies
Microsoft’s deployment overview distinguishes two implementation paths.
In connected environments, teams deploy Foundry Local as an Azure Arc extension and obtain prerequisites from online sources. In disconnected Azure Local environments, dependencies arrive through expansion packs and are installed from local registries.
The difference reaches beyond where model files are stored. Microsoft’s disconnected-environment documentation describes changes to extension installation, networking dependencies, certificate management, identity, telemetry and model sourcing.
Expansion packs populate the local edgeartifacts registry with required components. Catalog model artifacts are sourced locally, while bring-your-own models can come from a customer-managed OCI-compatible registry inside the disconnected environment. Networking dependencies that connected deployments obtain online are also imported through the packs.
For certificates, Microsoft says the azure-cert-manager extension is unavailable in disconnected environments. Teams instead install cert-manager and trust-manager using packaged assets.
Connected deployments typically use Microsoft Entra ID authentication. Disconnected deployments integrate with the local Active Directory infrastructure rather than public Entra ID endpoints. Kubernetes service account tokens are also available for in-cluster pod-to-pod inference in both environments.
Authorization needs particular attention. Microsoft documents separate connected roles for inference-only users and model-management users. In the disconnected path, its documented Contributor role is required for both write operations and inference. Teams should review those permissions against their intended separation of duties before copying an access design between environments.
Microsoft also says telemetry is not transmitted to it in disconnected deployments and that model evaluation data stays on the cluster. That supports a specific data-flow requirement, but does not establish regulatory compliance for the entire application.
An offline architecture still needs a controlled process for importing software and models, maintaining certificates and collecting diagnostics. Disconnected operation removes reliance on online sources during the documented deployment steps; lifecycle-management work remains.
NVIDIA’s Role Does Not Replace Workload-Level Controls
Microsoft describes NVIDIA accelerated computing as supporting demanding AI workloads and NVIDIA Confidential Computing as helping protect sensitive data, models and workloads during processing.
Those capabilities address important parts of the stack, but should not be interpreted as a complete sovereignty guarantee. Protection during processing does not, by itself, decide who may deploy a model, which information an agent may retrieve, or where application logs are retained.
The announcement also references NVIDIA AI Enterprise, its model ecosystem and RTX PRO infrastructure as options across software, models and compute. These references explain the collaboration’s breadth. The paper’s publication does not establish new availability, performance results or compliance approvals for those products.
Architects need to ask which capability addresses a documented threat or operational requirement. Hardware protection, identity controls and deployment authority solve different problems and need to be assessed separately.
A Pilot Should Test Failure and Model Replacement
Microsoft’s deployment guidance moves from exploration through proof of concept, pilot and production planning. It recommends evaluating model quality, response times, resource usage, authentication, monitoring, updates and operational ownership at the relevant stages.
For a regulated workload, two tests deserve special emphasis.
First, exercise the application under its expected connectivity failure conditions. Verify that necessary identity, retrieval, inference and certificate services remain available, not merely that the model process stays running. For example, a local assistant still needs access to its permitted document store if document retrieval is essential to answering requests.
Second, replace the model in a representative workflow. Check response quality, hardware requirements, endpoint behavior and access controls. A catalog containing several models demonstrates availability of options; a successful replacement demonstrates usable choice.
The pilot should also record who owns model approval, artifact imports, updates and recovery. Those responsibilities are part of deployment control, particularly where disconnected infrastructure cannot depend on routine online installation paths.
The Microsoft-NVIDIA guidance is most actionable when its questions become acceptance criteria. Preview-stage Foundry Local offers one documented way to investigate them, but local endpoints alone are not proof that a workload meets its sovereignty obligations.
Frequently Asked Questions
4 questions
1Did Microsoft and NVIDIA launch a new sovereign AI service?
No. Microsoft announced a Sovereign AI white paper developed with contributions from NVIDIA on October 5, 2026. It presents architecture guidance organized around control, choice, flexibility and resilience. References to Microsoft and NVIDIA products explain implementation options; the announcement is not a new model release or a newly shipped cloud service.
2Is Foundry Local on Azure Local generally available?
Foundry Local on Azure Local is available in preview, according to Microsoft’s documentation. It supports organizational inference on Azure Local infrastructure using Arc-enabled Kubernetes. Microsoft warns that preview features, approaches and processes may change or have limited capabilities before general availability. Teams should account for that during deployment planning.
3Can Foundry Local on Azure Local run without internet access?
Yes. Microsoft documents a disconnected deployment path that imports required components and model artifacts through expansion packs and local registries. Authentication integrates with local Active Directory rather than public Entra ID endpoints. Teams still need to plan certificate management, artifact imports, diagnostics and the application’s other dependencies.
4Does OpenAI compatibility make models interchangeable?
No. OpenAI-compatible request patterns help applications integrate with Foundry Local, but do not guarantee identical model behavior or output quality. Replacing a model still requires checking workload-specific quality, latency, capacity and operational controls. API compatibility is one part of portability, not proof that a replacement will meet the same requirements.
Sources
- October 5, 2026 announcementmicrosoft.com
- Foundry Local on Azure Local overviewlearn.microsoft.com
- deployment overviewlearn.microsoft.com
- disconnected-environment documentationlearn.microsoft.com





