AIbaseUpdated

Microsoft released the MAI Code1.1Flash model, and GitHub Copilot will support hybrid local and cloud-based reasoning.

On October 7, 2026, at the Windows and Surface press event, Microsoft announced that by the end of this month, local AI model support will be introduced for GitHub Copilot. Developers will be able to freely switch…

Microsoft announced at the Windows and Surface press conference on October 7, 2026, that it will introduce local AI model support for GitHub Copilot by the end of this month. This will allow developers to freely switch between cloud and device models or implement automatic scheduling.

To address the memory bottleneck and long context resource usage issues in edge inference, Microsoft has introduced an efficient MAI Code1.1Flash Mixture of Experts (MoE) model. This model has a total of 137 billion parameters and 6.8 billion activated parameters. By combining quantization and speculative decoding techniques, it significantly reduces memory usage while improving the encoding response speed on the device side. It is also the first to be integrated into the new Surface Laptop Ultra equipped with NVIDIA RTX Spark hardware.

In specific applications, users can choose the reasoning mode independently through the GitHub Copilot CLI, Copilot app, and Visual Studio Code. They can rely on Copilot’s backend to automatically coordinate resources, or specify calling local models such as MAI Code1.1Flash via the Windows ML provider or a compatible local endpoint that supports OpenAI. In terms of security, the system integrates Microsoft Execution Containers (MXC) built by the Windows team, utilizing the native container isolation mechanisms of each operating system to ensure the safety of code execution.

This upgrade indicates that the AI coding assistant is accelerating its evolution from a purely cloud-dependent model to a hybrid architecture combining the cloud and on-device capabilities. This move not only simplifies the deployment of local models, but also provides more efficient solutions for high-privacy, low-latency development scenarios.

Original source

AIbase

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original