AIbaseUpdated

Cloudflare releases open-weight decision model Clef-omni, adding audio and video multimodal input support

Cloudflare released the open-weight decision model Clef-omni on October 9. Building on its existing text and image processing capabilities, it adds support for audio and full video, and can process multimodal content…

Cloudflare officially announced the launch of the new open-weight decision model Clef-omni on October 9. Building on its existing text and image processing capabilities, the model further extends support for audio and full video formats, and can process multimodal content simultaneously through a single API call.

Multimodal upgrade and architecture analysis

In terms of multimodal support, the traditional Clef model mainly processed text, images, and static consecutive frames sampled at time intervals. The newly launched Clef-omni natively supports audio and video formats such as wav, mp3, mp4, and webm. This improvement eliminates the cumbersome steps developers previously needed to deploy additional speech transcription and audio/video splitting pipelines, allowing diverse input requirements to be handled efficiently with a single model.

According to official technical details, Clef-omni is built on Qwen3-Omni-30B-A3B-Instruct and retains its core understanding capabilities. The model is mainly aimed at structured decision tasks and does not generate regular text output. In terms of performance, the median response time for pure text requests is about 130 milliseconds, and about 150 milliseconds for images; processing a 21-second video with sound takes only about 1.5 seconds to complete scoring.

Price reduction and series model comparison

Alongside the release of the new model, Cloudflare also lowered the usage price of Clef-flash, with the cost per million input tokens cut from 0.09 USD to 0.038 USD, a reduction of about 58%, while the context window of its hosted version was also adjusted from 64k to 24k. Overall, the current tiered pricing of this model series meets the cost and performance needs of different developers.


Original source

AIbase

Content notes

Original publication and rights belong to the source.

Machine translation · Refer to the original