VYPR
researchPublished Oct 9, 2026· 1 source

Cloudflare Unveils Clef-omni, a Multimodal Decision Model, and Enhances Clef Family

Cloudflare has launched Clef-omni, an open-weight decision model capable of processing audio, video, image, and text inputs simultaneously, alongside price reductions and performance improvements for its Clef family of models.

Cloudflare continues its rapid innovation in the decision model space with the introduction of Clef-omni, a powerful new open-weight model designed to handle multimodal inputs. This release builds upon the recent launch of Clef and Clef-flash, demonstrating Cloudflare's commitment to iterating quickly and delivering advanced AI tools to the community. Clef-omni's key differentiator is its ability to process audio, video, image, and text data concurrently, moving beyond the text-centric limitations of many existing decision models.

Unlike previous models that might require separate pipelines for transcription or media stream separation, Clef-omni integrates these functionalities into a single, unified analysis. This approach aims to mimic human perception more closely, allowing for decisions based on a holistic understanding of various data types. The model is built on a Qwen3-Omni-30B-A3B-Instruct foundation, leveraging its inherent capability to process diverse media formats directly.

Technically, Clef-omni bypasses the overhead associated with transcription services by directly scoring all modalities within a single API call. Media elements are mapped into a unified sequence, synchronizing audio and video with visual frames for joint processing. The model employs a two-stage attention routing mechanism to gather evidence from all input types before cross-attending across the full context to compute confidence scores. A built-in lexical grammar ensures schema-constrained scoring, making it robust against variations in data structure and prompt formats.

Cloudflare highlights the speed and efficiency of Clef-omni, with text-only decisions returning in approximately 130 milliseconds, image inputs in about 150 milliseconds, and even a 21-second video clip with audio being scored in about 1.5 seconds. This performance is achieved through optimizations that skip output token generation, as Clef models are designed for decision-making rather than generative tasks.

In addition to Clef-omni, Cloudflare has also made its Clef-flash model more accessible by significantly reducing its price. Clef-flash is now available at $0.038 per M input tokens, making it cheaper than the competing Jev model and incentivizing broader adoption. The original Clef model remains priced at $0.24 per M input tokens, while Clef-omni is launched at $0.15 per M input tokens.

Cloudflare emphasizes that these releases are just the beginning, reflecting their DNA of innovation and iteration. The company aims to continuously improve its models in terms of speed, cost, and capabilities, pushing the boundaries of what open-weight decision models can achieve. The models are available as open-weights on HuggingFace, with detailed developer documentation provided.

Performance benchmarks shared by Cloudflare show Clef-omni performing competitively across various evaluation datasets, including BFCL, ToolRet, and API-Bank, often matching or exceeding the performance of Clef and Clef-flash, and significantly outperforming Jev in several metrics. This demonstrates the model's effectiveness in handling complex, multimodal decision-making tasks.

The release of Clef-omni and the pricing adjustments for Clef-flash underscore Cloudflare's strategy to democratize access to advanced AI technologies. By offering powerful, open-weight multimodal models at competitive price points, Cloudflare aims to empower developers and organizations to build more sophisticated and context-aware applications.

Synthesized by Vypr AI