Cloudflare's Clef Evolves: Multimodal, Faster, Cheaper

Alps Wang

Alps Wang

Oct 10, 2026 · 1 views

Unpacking Cloudflare's Multimodal AI Leap

Cloudflare's rapid iteration on its Clef family of models, particularly with the introduction of Clef-omni, marks a significant step towards more integrated and intuitive AI decision-making. The native handling of audio, video, image, and text in a single model bypasses complex, resource-intensive pipelines, democratizing advanced AI capabilities. The aggressive pricing strategy for Clef-flash, now cheaper than Jev, and the performance gains in Clef, further enhance accessibility and adoption. The open-weight nature of these models fosters community innovation and allows for deeper integration into developer workflows, potentially accelerating the development of sophisticated AI agents and applications. The clear emphasis on schema-constrained, calibrated decisions makes these models particularly suitable for practical, real-world use cases where accuracy and reliability are paramount.

However, while the ambition is commendable, the long-term implications of a reduced context window for Clef-flash warrant careful consideration. Although Cloudflare cites usage data to justify this, it could limit certain advanced applications that rely on extensive historical context. Furthermore, the underlying foundation of Clef-omni, based on a Qwen3-Omni-30B-A3B-Instruct MoE, while powerful, implies a certain level of complexity and resource requirement for self-hosting, even if Cloudflare offers it as a managed service. The benchmarks, while impressive, are internal and against specific datasets; real-world performance across a wider array of diverse and noisy inputs will be the ultimate test. The comparison with Jev is valuable, but the broader competitive landscape of multimodal models from other major players will continue to evolve rapidly, demanding continuous innovation from Cloudflare.

Key Points

  • Cloudflare has launched Clef-omni, a multimodal decision model capable of processing audio, video, image, and text inputs simultaneously.
  • This eliminates the need for complex, cascading model pipelines for cross-modal analysis.
  • Clef-flash has been made significantly cheaper, priced at $0.038 per M input tokens, making it more affordable than Jev.
  • The Clef model has also seen performance improvements, with faster inference speeds achieved through serving infrastructure optimizations.
  • Clef-omni is built on a Qwen3-Omni-30B-A3B-Instruct MoE foundation and utilizes a novel two-stage attention routing mechanism for efficient multimodal processing.
  • A trade-off for Clef-flash's lower price is a reduced context window of 24k tokens (for hosted version), while the model weights remain capable of 256k context for self-hosting.
  • The models are open-weight and available on HuggingFace, promoting community development and integration.

Article Image


📖 Source: Introducing Clef-omni with full multimodality, plus a faster Clef and a cheaper Clef-flash

Related Articles

Comments (0)

No comments yet. Be the first to comment!