IT Home On October 9, the tech media outlet NeoWin posted a blog post yesterday (October 8) stating that Microsoft is collaborating with the llama.cpp and GGUF open-source communities to upgrade Windows ML in the Windows 11 system. It will experientially support reasoning of GGUF and ONNX format models natively, and introduce the Windows ML Runtime API.
During the Windows and Surface October event, Microsoft introduced the latest developments in Windows ML. Windows ML is a local AI inference framework developed by Microsoft for Windows 11. Developers can run AI models on devices, build applications such as text generation and speech recognition, and reduce their dependence on hardware-specific optimizations.

Microsoft says through community collaboration, developers can now import GGUF models from Hugging Face using the experimental integration of llama.cpp, and run them directly on their local machines. In terms of performance, Microsoft has also worked with NVIDIA to contribute code to llama.cpp and related performance optimizations.

Windows ML introduces an experimental native inference path for Windows, which can run both ONNX and GGUF models. The official name of this path is Windows ML Runtime API; Microsoft claims it offers higher performance, deeper system integration, and finer-grained control capabilities.
IT Home note: GGUF is a model file format commonly used for running large language models locally. Developers can obtain GGUF models from platforms such as Hugging Face, and load and infer them on a personal computer using tools like llama.cpp.
And ONNX is the Open Neural Network Exchange format, used for migrating AI models between different frameworks and hardware. ONNX Runtime is its common runtime environment, which supports the inference deployment of models on Windows devices.

With the support of the Windows ML Runtime API, Microsoft has also introduced the Text Generation API and Speech Recognition API. Developers can use these two interfaces to build text generation and speech recognition features.

Microsoft said that the ONNX Runtime API is still fully supported; future optimizations for Windows will be prioritized for the Windows ML Runtime API. Both runtime technologies are included in the current version, and developers can choose the timing of migration based on their familiarity with them.
