Cloudflare recently officially released the new open-weight decision-making model, Clef-omni. This model has been significantly upgraded based on the existing Clef series. It not only natively supports text and images, but also for the first time offers direct processing capabilities for audio and full video inputs. Currently, its model weights have been fully open-sourced on the Hugging Face platform.
In terms of technical implementation and functional features, Clef-omni has completely transformed traditional processing pipelines. Previous versions of Clef often required breaking down video inputs into continuous static images based on the time axis before processing; however, the new Clef-omni directly supports mainstream audio and video formats such as wav, mp3, mp4, and webm. This means developers no longer need to build complex speech transcription and audio/video splitting pipelines separately; they can achieve unified processing of text, images, audio, and video with a single API call.
In terms of performance and benchmark results, this model is built based on the Qwen3-Omni-30B-A3B-Instruct base architecture, and its core understanding capabilities are fully preserved. As a specialized model focused on structured decision-making tasks, it does not produce regular long texts. According to the officially published actual test data, the median response time for pure text requests is approximately 130 milliseconds, while for image requests it is about 150 milliseconds; moreover, processing a video with sound that lasts up to 21 seconds can also be efficiently completed in just about 1.5 seconds to score it.
With the introduction of the new model, Cloudflare has also made price adjustments and optimizations to other product lines. Specifically, the input price for Clef-flash has seen a significant reduction of approximately 58%, dropping from $0.09 per million input tokens to $0.038, while the hosting version’s context window has been adjusted from 64k to 24k. Overall, with the open-source nature of Clef-omni and the enhancement of multimodal processing capabilities, it provides a more efficient and cost-effective underlying support for global developers in building complex intelligent workflows and automated decision systems.