innFactory AI Consulting from Rosenheim adds Clef because Cloudflare extends the category of System One models we explained in Jev by TypeSafe: the AI model that writes no text with an open, multimodal and hosted offering. This page is based on the announcement on the Cloudflare blog, the Workers AI changelog, the model pages for Clef and Clef-flash, the Workers AI price list and the model cards on Hugging Face. As of 2 October 2026.
What is Clef?
On 1 October 2026 Cloudflare released Clef and Clef-flash, the first models the company has trained itself. Both are decision models: they take a state (a support case, a web page, an agent trace, an image) plus a list of typed questions and return probabilities over predefined answers for each question. They generate no text. The question types are the same as with Jev: noul (yes/no as a probability), choice (one option from a list with a probability per option) and score (a level on a rubric). Cloudflare describes the models as fully compatible with the Jev API.
| Clef | Clef-flash | |
|---|---|---|
| Workers AI ID | @cf/cloudflare/clef | @cf/cloudflare/clef-flash |
| Base | Qwen3.8-27B (frozen) with vision encoder | Qwen3.5-9B (frozen) with vision encoder |
| Parameters | 27 billion | 9 billion |
| Context (Workers AI) | 65,536 tokens | 65,536 tokens |
| Price (Workers AI) | USD 0.24 per 1M input tokens | USD 0.09 per 1M input tokens |
| Median latency (Cloudflare, H200) | 209.3 ms | 38.8 ms |
| Licence | Apache 2.0 | Apache 2.0 |
Technically, a joint schema head sits on top of the frozen Qwen backbone: a small transformer that routes the evidence in the state to the questions and emits exactly one logit per option. Per the blog it was trained with label-smoothed cross-entropy, a Brier loss for calibration and Reinforcement Learning for Calibrated Decisions (RLCD) as a secondary target.
Inputs and limits on Workers AI
- State: text, JSON objects or arrays, up to 4 images (PNG, JPEG, WebP, at most 4 MiB and 16 megapixels each) and video; request body at most 13 MiB, decoded image data at most 8 MiB.
- Questions: 1 to 64 per request, question IDs up to 100 characters.
- Response: model identifier,
answerskeyed like the questions, andusage. - Call: Workers binding
env.AI.run(), REST endpoint/ai/runor via the Cloudflare AI Gateway; examples for TypeScript, Python and cURL in the docs. - Billing: Workers AI bills in neurons (Clef 21,818, Clef-flash 8,182 neurons per 1 million input tokens); 10,000 neurons per day are free, beyond that USD 0.011 per 1,000 neurons on the Workers Paid plan. No output price is listed.
Benchmarks: Cloudflare’s figures
Cloudflare publishes a table against Jev in the blog; per the changelog Clef leads on 7 of 10 decision benchmarks. These are vendor figures.
| Benchmark | Clef | Clef-flash | Jev (per Cloudflare) |
|---|---|---|---|
| BFCL, case exact | 98.47 | 98.76 | 95.75 |
| ToolRet, nDCG@10 | 69.19 | 66.43 | 65.28 |
| API-Bank, accuracy | 91.93 | 93.11 | 88.19 |
| BANKING77, macro-F1 | 94.20 | 90.93 | 79.74 |
| CLINC150+OOS, macro-F1 | 97.43 | 66.77 | 89.27 |
| GPQA Diamond (knowledge) | 48.0 | 51.0 | 78.3 |
| MMLU-Pro (knowledge) | 65.9 | 65.3 | 82.7 |
Two things stand out. First, Jev leads clearly on the knowledge benchmarks, Clef on the tool and intent tasks. Second, Clef-flash collapses on CLINC150 with out-of-scope detection (66.77 versus 97.43 for Clef). Anyone picking Clef-flash for its latency should test out-of-scope cases separately.
In the workflow evaluations published by TypeSafe (invoice processing, customer service, security incidents) all three models are close per Cloudflare: 64.7 / 57.1 / 61.8 on invoices, 76.3 / 77.0 / 76.0 in customer service, 62.9 / 61.7 / 61.7 on incidents (Clef / Clef-flash / Jev).
Latency (Cloudflare measurement on an H200): Clef 209.3 ms median and 238.6 ms p95, Clef-flash 38.8 ms and 122.4 ms, Jev 524.1 ms and 536.0 ms. In the same table Cloudflare lists Laya at 5.8 ms median but 222.5 ms p95.
Independent single test: a developer compared Clef-flash locally on an RTX 3090 against the Jev API on 42 decisions of an autonomous agent: Jev 71.4%, Clef-flash 66.7%, the same decision in 83% of cases; identical on computer-use actions and supervision, 10 instead of 12 of 20 hits on triage. The test is small and not peer-reviewed, but it shows the order of magnitude. A practical note from it: community GGUF quantisations lack the decision head, PyTorch is needed for full functionality; Clef (27B) needs around 54 GB of GPU memory in BF16.
Data protection, regions and self-hosting
For the hosted variant Cloudflare states that requests and responses are neither read, stored nor used for training, unless the customer uses the fine-tuning offering. The Workers AI data usage docs confirm that customer content is not used to train models or improve services. Workers AI runs on Cloudflare’s global network; the Data Localization Suite documentation does not list Workers AI as a product with region pinning (as of 2 October 2026). Anyone who needs a contractual commitment to processing in the EU should clarify that with Cloudflare or take the second route.
The second route is self-hosting: the weights of both models are on Hugging Face under Apache 2.0 (Cloudflare/clef, Cloudflare/clef-flash). The model card describes operation with Transformers, vLLM and SGLang (OpenAI-compatible APIs), via Docker, and a systemone() function for the Jev-compatible interface; Cloudflare tested on a single H200 in BF16. That lets you run Clef on your own infrastructure or with an EU cloud provider, with full data control.
RL fine-tuning platform
Together with Clef, Cloudflare announced a platform for reinforcement learning fine-tuning. The building blocks: AI Gateway captures production data, Workers AI generates rollouts, Cloudflare Containers serve as the RL sandbox, a new Trainer updates the weights, and the result is redeployed via Workers AI with a bring-your-own model. At launch it is a guided offering with Cloudflare’s forward-deployed engineers via a design-partner form; a self-serve version is announced, pricing is not published.
Positioning: Clef, Jev and Laya
| Clef / Clef-flash | Jev | Laya | |
|---|---|---|---|
| Licence | Apache 2.0, open weights | proprietary API | Apache 2.0, open weights |
| Operation | Workers AI or self-hosting | API only (US) | self-hosting only |
| Size | 27B / 9B (Qwen base) | not published | 0.3 to 0.4B (encoder) |
| Inputs | text, JSON, images, video | text, JSON | text, JSON |
| Context | 65,536 tokens | 64,000 tokens combined | 512 to 1,024 tokens, windows up to 8,192 |
| Hosted price | USD 0.24 / 0.09 per 1M input | USD 0.042 per 1M input | no hosted API |
| EU residency | via self-hosting | not documented | via self-hosting |
Clef closes the gap between Jev (strong, but closed and US-hosted) and Laya (open, but small): an open model with an LLM-sized backbone, image and video input and a hosted option with a price. The price is above Jev, the latency of Clef-flash below. None of the three replaces an LLM; for text, summaries and code, models such as Qwen, Mistral or GPT-OSS remain in charge.
Our recommendation
For companies in the DACH region Clef is interesting above all as a self-hosted decision model: open licence, Jev-compatible interface, image and video input, and with Clef-flash a variant that runs on a single GPU. Check three points before adopting it: out-of-scope detection with Clef-flash, language quality on German data (Cloudflare publishes English benchmarks only) and, for hosted operation, the contractual situation regarding the processing region. We assess its use as a front stage to language models in CompanyGPT or as a router in the AI Gateway per project. For an assessment of whether a decision model fits your architecture, contact innFactory AI Consulting.
