Clef-flash
Smaller, faster sibling of Clef (Qwen3.5-9B backbone). Self-hosting needs ~41 GB VRAM at 64k context (The Register). Workers AI price not verified.
- Input
- text, image, video
- Output
- text
- License
- apache-2.0
- Verified
- 2026-10-02
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Cloudflare Workers AI | @cf/cloudflare/clef-flash | — | docs |
| Hugging Face | — | huggingface.co/Cloudflare/clef-flash | — |
Notable capabilities (1)
- Low-latency decision model: Qwen3.5-9B-based option scorer; Cloudflare reports 38.8 ms median and 122.4 ms p95 latency on its 43-benchmark set and top scores on BFCL (98.76) and API-Bank (93.11). source
Fast open-weights decision model from Cloudflare. Same Jev-compatible API as cloudflare-clef. Hugging Face repo name taken from Cloudflare's blog.
Timeline entry
- Cloudflare releases Clef and Clef-flash, open-weight multimodal decision models that are Jev-API compatible ★★★
On Oct 1, 2026 Cloudflare released Clef (built on Qwen3.8-27B) and Clef-flash (built on Qwen3.5-9B), Apache-2.0 open-weight "decision models". Like TypeSafe's Jev they return probabilities for typed options instead of generating text, but they also accept images and video and have a 64k context…