Windows ML can now accept GGUF language models through an experimental llama.cpp integration, giving developers a local inference path without converting those models to ONNX. The release also exposes an OpenAI-compatible endpoint, so applications already using the OpenAI SDK can try a Windows-hosted model by changing their connection settings.
In its October 7, 2026 announcement, Microsoft describes two new task-specific APIs: Text Generation for GGUF and ONNX language models, and Speech Recognition for an ONNX Whisper model. Beneath them sits a new experimental Windows-native Runtime API for developers who need more control over model execution and multi-model pipelines.
A shared application-facing layer over different model formats and execution engines could reduce integration work for Windows developers evaluating local AI. Microsoft labels both the llama.cpp integration and native runtime experimental, however. Teams should distinguish that preview from the existing supported ONNX Runtime path and the separate security and hardware announcements surrounding it.
GGUF Joins ONNX Rather Than Replacing It
The Text Generation API accepts language models in either GGUF or ONNX format. Microsoft says Windows ML selects the appropriate execution engine for the supplied model, including llama.cpp for GGUF. Developers can therefore bring a GGUF model into the Windows ML stack without first turning it into an ONNX deployment. Microsoft specifically presents downloading a GGUF model from Hugging Face and running it locally as a use case.
With automatic engine selection, an application's text-generation code can sit above the format-specific choice. A developer can evaluate a GGUF model through the same higher-level API surface used for an ONNX language model, without building a separate application integration around each engine.
The formats remain distinct. Accepting both through one API doesn't make their underlying execution paths interchangeable, and the announcement doesn't establish that every model, quantization variant, or device combination will work.
Existing ONNX applications don't need to migrate immediately: Microsoft says the familiar ONNX Runtime APIs remain fully supported and ship alongside the experimental Windows-native Runtime API.
For a team with a working ONNX deployment, the update creates an additional evaluation path. By itself, it gives no reason to replace a known-good inference integration. The most useful first experiment is whether a particular GGUF model improves the application's output quality, memory requirements, or responsiveness on its target PCs.
The OpenAI-Compatible Endpoint Reduces Client Rework
Microsoft's example starts a Windows ML server with a GGUF model, then points the OpenAI SDK at its OpenAI-compatible local endpoint.
The published server command is:
WinMLServer.exe model.gguf --model-id qwen2.5-0.5b --target gpu --port 8080
The client then uses the local base URL and the access key printed when the server starts:
from openai import OpenAI
client = OpenAI(
base_url="http://127.0.0.1:8080/v1",
api_key="<access key printed on startup>",
)
stream = client.chat.completions.create(
model="qwen2.5-0.5b",
messages=[
{
"role": "user",
"content": "Summarize the benefits of local model inference.",
}
],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
This follows Microsoft's published example, with a different prompt. It illustrates the client connection; it isn't a complete installation or deployment procedure.
Developers can retain the familiar SDK and chat-completions request pattern while directing inference to a local server. That can make an existing application a useful test harness for comparing a cloud workflow with an on-device model.
There are limits to what this compatibility establishes. The local model isn't an OpenAI model, and the example doesn't demonstrate support for every OpenAI API feature. It demonstrates streamed chat completions. Teams should separately check any features their application depends on, such as tool calling or structured responses.
The model identifier refers to the model served locally. Changing the endpoint doesn't preserve the behavior of a previously used cloud model: prompt adherence, response quality, latency, and resource use still need evaluation against the actual GGUF model.
Text and Speech Can Share a Local Pipeline
The first task-specific release covers two workloads:
- Text Generation: Runs a developer-supplied GGUF or ONNX language model.
- Speech Recognition: Transcribes audio using a developer-supplied ONNX Whisper model.
Microsoft describes combining them by transcribing voice input and passing the resulting text to a GGUF language model. This gives developers a concrete starting point for a local voice interface without requiring both stages to use the same model format.
Speech recognition produces the text input; the language model handles the subsequent generation task. Developers still supply the models and decide what the application does with their outputs.
The release is limited to these first task-specific APIs. The announcement doesn't establish equivalent high-level APIs for image generation or every other model category.
A sensible prototype would begin with one of those supported tasks, then measure the complete workflow. A fast language model won't necessarily make a voice application responsive if transcription or model loading dominates the delay.
The Native Runtime Adds Explicit Device Placement
The experimental Windows-native Runtime API underlies the higher-level text and speech APIs and gives developers finer control over execution.
Microsoft says it supports composing multi-model pipelines with explicit device placement at each stage across CPU, GPU, and NPU. An application can specify where individual stages should execute instead of treating the workflow as one undifferentiated inference request.

Engine selection and device placement are separate decisions. Choosing llama.cpp for a GGUF model identifies the execution engine; assigning a pipeline stage to a particular device concerns hardware placement. The announcement should not be read as a guarantee that every GGUF model can run on every Windows NPU.
Microsoft also describes two other capabilities:
- Windows-native inputs, including images, video frames, audio buffers, and text, through what it calls efficient zero-copy paths.
- Ahead-of-time model loading and compilation into a ready-to-run artifact, with device and execution-policy choices applied.
Input conversion, memory movement, and startup can affect the user experience even when generation itself is fast. These features could help when model execution accounts for only part of an application's cost.
The supplied announcement describes these capabilities without independently measured results. Microsoft also calls pipeline execution reproducible and predictable; developers shouldn't extend that statement into a promise that generative text outputs will always be identical.
Performance Claims Need Hardware-Specific Evidence
Microsoft says its llama.cpp contributions, made with NVIDIA and the broader community, include CUDA kernel optimization, kernel fusion, improved CPU-GPU scheduling, weight repacking, and CUDA graphs. It also lists additions involving speculative decoding, multi-GPU execution, and other capabilities.
Sources
- October 7, 2026 announcementdevblogs.microsoft.com





