↓ Skip to main content
9 – 17 UHR +49 8031 3508270 LUITPOLDSTR. 9, 83022 ROSENHEIM
DE / EN
SPECIALIZED Cloudflare USA

Cloudflare Clef (Clef & Clef-flash)

Clef and Clef-flash are Cloudflare's first in-house trained models: open, multimodal decision models (Apache 2.0, 27B and 9B parameters on Qwen bases) that answer typed questions with probabilities instead of text and are compatible with the Jev API. On Workers AI at USD 0.24 and 0.09 per 1 million input tokens, weights on Hugging Face for self-hosting in the EU. As of 2 October 2026.

License Apache 2.0
GDPR Hosting Available
Context 65536 Tokens
Modality Text, JSON, Image, Video → Typed decisions (choice, score, noul) with probabilities

Versions

Overview of available model variants

ModelReleaseEUStrengthsWeaknessesStatus
Clef (@cf/cloudflare/clef, 27B)
1 October 2026
Cloudflare's first in-house trained model: 27B parameters, post-trained on a frozen Qwen3.8-27B with vision encoder Multimodal state: text, JSON, up to 4 images (PNG, JPEG, WebP, 4 MiB and 16 megapixels each) and video; 65,536 token context, 1 to 64 questions per request Per Cloudflare benchmarks: BFCL 98.47, ToolRet nDCG@10 69.19, API-Bank 91.93, BANKING77 94.20 macro-F1, CLINC150+OOS 97.43 macro-F1 (Jev: 95.75 / 65.28 / 88.19 / 79.74 / 89.27) Jev API compatible (question types noul, choice, score, same response format); non-autoregressive, one logit per option USD 0.24 per 1 million input tokens on Workers AI, callable via Workers binding, REST /ai/run or Cloudflare AI Gateway Apache 2.0 on Hugging Face (Cloudflare/clef); self-hosting with Transformers, vLLM, SGLang (OpenAI-compatible), Docker and a systemone() function Trained with label-smoothed cross-entropy, Brier loss for calibration and RLCD as a secondary target (per the blog)
Median latency 209.3 ms (p95 238.6 ms) on an H200 per Cloudflare, roughly five times slower than Clef-flash Knowledge benchmarks below Jev: GPQA Diamond 48.0 versus 78.3, MMLU-Pro 65.9 versus 82.7 (Cloudflare table) Self-hosting needs around 54 GB of GPU memory in BF16; community GGUF quantisations lack the decision head, PyTorch required (independent test report) All comparison figures come from Cloudflare; the only independent evaluation so far is a single test with 42 decisions No documented region pinning for Workers AI inference; language coverage not stated in the model card
Current
Clef-flash (@cf/cloudflare/clef-flash, 9B)
1 October 2026
9B parameters based on Qwen3.5-9B with vision encoder; median latency 38.8 ms, p95 122.4 ms on an H200 (Cloudflare), roughly 13 times faster than Jev in the same measurement USD 0.09 per 1 million input tokens on Workers AI; same inputs and limits as Clef (65,536 tokens, 64 questions, 4 images, video) Per Cloudflare on par with or ahead of Clef on tool benchmarks: BFCL 98.76, API-Bank 93.11, BANKING77 90.93 Independent single test (42 agent decisions): 66.7% versus 71.4% for Jev, 83% agreement with the API, runs on an RTX 3090 Jev API compatible, Apache 2.0, self-hosting with far lower memory needs than Clef
CLINC150+OOS only 66.77 macro-F1 (Clef 97.43, Jev 89.27): out-of-scope detection markedly weaker (Cloudflare table) GSM8K 67.3 versus 80.8 for Clef per the model card; triage tasks in the independent test show the largest gap to Jev (10 of 20 versus 12 of 20) Model card defaults to 16,384 tokens of input length when self-hosting; 65,536 applies on Workers AI Comparison figures from the vendor, language coverage not stated
Current

Use Cases

Typical applications for this model

Ticket and support triage with confidence
Classification of web pages, domains and bots (Cloudflare examples)
Trust and safety review of submissions
Decision nodes and supervision in AI agents
Image and video classification with typed answers
Routing in front of expensive reasoning models

Technical Details

API, features and capabilities

API & Availability
Availability Public on Cloudflare Workers AI (binding env.AI.run, REST /ai/run, AI Gateway); weights on Hugging Face for self-hosting (Transformers, vLLM, SGLang, Docker)
Latency (TTFT) 209.3 ms (Clef) or 38.8 ms (Clef-flash) median on an H200 (Cloudflare measurement)
Features & Capabilities
Structured Output Vision
Training & Knowledge
Knowledge Cutoff not documented (backbones Qwen3.8-27B and Qwen3.5-9B)
Fine-Tuning Available (Cloudflare RL fine-tuning platform (guided by forward-deployed engineers, self-serve announced), Your own fine-tuning of the open weights)
Language Support
Best Quality English
Supported Not stated; the Qwen backbones are multilingual, Cloudflare publishes no language evaluation
Cloudflare's benchmarks are in English; evaluate on your own German data before deployment

Hosting & Compliance

GDPR-compliant hosting options and licensing

GDPR-Compliant Hosting Options
Self-hosted
Your own EU infrastructure or EU cloud
Recommended for EU residency: Apache 2.0 weights, vLLM/SGLang/Transformers, Clef around 54 GB of GPU memory in BF16
Cloudflare Workers AI
Cloudflare global network
No reading, storing or training on customer requests per Cloudflare; no documented region pinning of inference in the Data Localization Suite (as of 2 October 2026)
License & Hosting
License Apache 2.0
Security Filters None (decision model without text generation)
Enterprise Support Yes
On-Premise

Benchmarks

Performance comparison with standardized tests

BFCL case exact (Clef, Cloudflare figure)
98.47
BANKING77 macro-F1 (Clef, Cloudflare figure)
94.2
CLINC150+OOS macro-F1 (Clef, Cloudflare figure)
97.43
API-Bank accuracy (Clef-flash, Cloudflare figure)
93.11
MMLU-Pro (Clef, Cloudflare figure)
65.9

innFactory AI Consulting from Rosenheim adds Clef because Cloudflare extends the category of System One models we explained in Jev by TypeSafe: the AI model that writes no text with an open, multimodal and hosted offering. This page is based on the announcement on the Cloudflare blog, the Workers AI changelog, the model pages for Clef and Clef-flash, the Workers AI price list and the model cards on Hugging Face. As of 2 October 2026.

What is Clef?

On 1 October 2026 Cloudflare released Clef and Clef-flash, the first models the company has trained itself. Both are decision models: they take a state (a support case, a web page, an agent trace, an image) plus a list of typed questions and return probabilities over predefined answers for each question. They generate no text. The question types are the same as with Jev: noul (yes/no as a probability), choice (one option from a list with a probability per option) and score (a level on a rubric). Cloudflare describes the models as fully compatible with the Jev API.

ClefClef-flash
Workers AI ID@cf/cloudflare/clef@cf/cloudflare/clef-flash
BaseQwen3.8-27B (frozen) with vision encoderQwen3.5-9B (frozen) with vision encoder
Parameters27 billion9 billion
Context (Workers AI)65,536 tokens65,536 tokens
Price (Workers AI)USD 0.24 per 1M input tokensUSD 0.09 per 1M input tokens
Median latency (Cloudflare, H200)209.3 ms38.8 ms
LicenceApache 2.0Apache 2.0

Technically, a joint schema head sits on top of the frozen Qwen backbone: a small transformer that routes the evidence in the state to the questions and emits exactly one logit per option. Per the blog it was trained with label-smoothed cross-entropy, a Brier loss for calibration and Reinforcement Learning for Calibrated Decisions (RLCD) as a secondary target.

Inputs and limits on Workers AI

  • State: text, JSON objects or arrays, up to 4 images (PNG, JPEG, WebP, at most 4 MiB and 16 megapixels each) and video; request body at most 13 MiB, decoded image data at most 8 MiB.
  • Questions: 1 to 64 per request, question IDs up to 100 characters.
  • Response: model identifier, answers keyed like the questions, and usage.
  • Call: Workers binding env.AI.run(), REST endpoint /ai/run or via the Cloudflare AI Gateway; examples for TypeScript, Python and cURL in the docs.
  • Billing: Workers AI bills in neurons (Clef 21,818, Clef-flash 8,182 neurons per 1 million input tokens); 10,000 neurons per day are free, beyond that USD 0.011 per 1,000 neurons on the Workers Paid plan. No output price is listed.

Benchmarks: Cloudflare’s figures

Cloudflare publishes a table against Jev in the blog; per the changelog Clef leads on 7 of 10 decision benchmarks. These are vendor figures.

BenchmarkClefClef-flashJev (per Cloudflare)
BFCL, case exact98.4798.7695.75
ToolRet, nDCG@1069.1966.4365.28
API-Bank, accuracy91.9393.1188.19
BANKING77, macro-F194.2090.9379.74
CLINC150+OOS, macro-F197.4366.7789.27
GPQA Diamond (knowledge)48.051.078.3
MMLU-Pro (knowledge)65.965.382.7

Two things stand out. First, Jev leads clearly on the knowledge benchmarks, Clef on the tool and intent tasks. Second, Clef-flash collapses on CLINC150 with out-of-scope detection (66.77 versus 97.43 for Clef). Anyone picking Clef-flash for its latency should test out-of-scope cases separately.

In the workflow evaluations published by TypeSafe (invoice processing, customer service, security incidents) all three models are close per Cloudflare: 64.7 / 57.1 / 61.8 on invoices, 76.3 / 77.0 / 76.0 in customer service, 62.9 / 61.7 / 61.7 on incidents (Clef / Clef-flash / Jev).

Latency (Cloudflare measurement on an H200): Clef 209.3 ms median and 238.6 ms p95, Clef-flash 38.8 ms and 122.4 ms, Jev 524.1 ms and 536.0 ms. In the same table Cloudflare lists Laya at 5.8 ms median but 222.5 ms p95.

Independent single test: a developer compared Clef-flash locally on an RTX 3090 against the Jev API on 42 decisions of an autonomous agent: Jev 71.4%, Clef-flash 66.7%, the same decision in 83% of cases; identical on computer-use actions and supervision, 10 instead of 12 of 20 hits on triage. The test is small and not peer-reviewed, but it shows the order of magnitude. A practical note from it: community GGUF quantisations lack the decision head, PyTorch is needed for full functionality; Clef (27B) needs around 54 GB of GPU memory in BF16.

Data protection, regions and self-hosting

For the hosted variant Cloudflare states that requests and responses are neither read, stored nor used for training, unless the customer uses the fine-tuning offering. The Workers AI data usage docs confirm that customer content is not used to train models or improve services. Workers AI runs on Cloudflare’s global network; the Data Localization Suite documentation does not list Workers AI as a product with region pinning (as of 2 October 2026). Anyone who needs a contractual commitment to processing in the EU should clarify that with Cloudflare or take the second route.

The second route is self-hosting: the weights of both models are on Hugging Face under Apache 2.0 (Cloudflare/clef, Cloudflare/clef-flash). The model card describes operation with Transformers, vLLM and SGLang (OpenAI-compatible APIs), via Docker, and a systemone() function for the Jev-compatible interface; Cloudflare tested on a single H200 in BF16. That lets you run Clef on your own infrastructure or with an EU cloud provider, with full data control.

RL fine-tuning platform

Together with Clef, Cloudflare announced a platform for reinforcement learning fine-tuning. The building blocks: AI Gateway captures production data, Workers AI generates rollouts, Cloudflare Containers serve as the RL sandbox, a new Trainer updates the weights, and the result is redeployed via Workers AI with a bring-your-own model. At launch it is a guided offering with Cloudflare’s forward-deployed engineers via a design-partner form; a self-serve version is announced, pricing is not published.

Positioning: Clef, Jev and Laya

Clef / Clef-flashJevLaya
LicenceApache 2.0, open weightsproprietary APIApache 2.0, open weights
OperationWorkers AI or self-hostingAPI only (US)self-hosting only
Size27B / 9B (Qwen base)not published0.3 to 0.4B (encoder)
Inputstext, JSON, images, videotext, JSONtext, JSON
Context65,536 tokens64,000 tokens combined512 to 1,024 tokens, windows up to 8,192
Hosted priceUSD 0.24 / 0.09 per 1M inputUSD 0.042 per 1M inputno hosted API
EU residencyvia self-hostingnot documentedvia self-hosting

Clef closes the gap between Jev (strong, but closed and US-hosted) and Laya (open, but small): an open model with an LLM-sized backbone, image and video input and a hosted option with a price. The price is above Jev, the latency of Clef-flash below. None of the three replaces an LLM; for text, summaries and code, models such as Qwen, Mistral or GPT-OSS remain in charge.

Our recommendation

For companies in the DACH region Clef is interesting above all as a self-hosted decision model: open licence, Jev-compatible interface, image and video input, and with Clef-flash a variant that runs on a single GPU. Check three points before adopting it: out-of-scope detection with Clef-flash, language quality on German data (Cloudflare publishes English benchmarks only) and, for hosted operation, the contractual situation regarding the processing region. We assess its use as a front stage to language models in CompanyGPT or as a router in the AI Gateway per project. For an assessment of whether a decision model fits your architecture, contact innFactory AI Consulting.

Cost estimation for this model

For up-to-date token pricing, model variants and EU availability, see our sister project ai-prices.eu. It helps you compare and estimate the operational cost of leading AI models for your specific use case.

Compare prices on ai-prices.eu

ai-prices.eu is a project by innFactory AI Consulting GmbH and provides transparent cost estimates for leading AI models.

Frequently Asked Questions

What is Cloudflare Clef?

Clef is a decision model released by Cloudflare on 1 October 2026, the first model Cloudflare trained itself. It takes a state (text, JSON, images or video) and a list of typed questions and returns probabilities over predefined answers for each question, without generating text. Clef has 27 billion parameters based on Qwen3.8-27B, Clef-flash 9 billion based on Qwen3.5-9B. Both are available under Apache 2.0 on Hugging Face and hosted on Cloudflare Workers AI.

What is the difference between Clef and Clef-flash?

Clef (27B) is the more accurate model, Clef-flash (9B) the faster one: on an H200 Cloudflare measures a median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash. On Workers AI Clef costs USD 0.24 and Clef-flash USD 0.09 per 1 million input tokens. In Cloudflare's benchmarks the two are close on tool and intent tasks; on CLINC150 with out-of-scope detection Clef-flash drops markedly to 66.77 versus 97.43 macro-F1 for Clef.

What does Clef cost on Workers AI?

The Workers AI price list (as of 2 October 2026) states USD 0.24 per 1 million input tokens for Clef (21,818 neurons) and USD 0.09 for Clef-flash (8,182 neurons); no output price is listed. Workers AI bills in neurons, 10,000 neurons per day are free, beyond that USD 0.011 per 1,000 neurons on the Workers Paid plan.

Can Clef be used in a GDPR-compliant way?

Two routes: hosted on Workers AI, Clef runs on Cloudflare's global network; Cloudflare states it does not read, store or train on requests and responses, but documents no region pinning for Workers AI inference in the Data Localization Suite (as of 2 October 2026). Anyone who must keep processing in the EU self-hosts the Apache 2.0 weights, for example via vLLM on their own infrastructure or with an EU cloud provider; Clef needs around 54 GB of GPU memory in BF16, Clef-flash considerably less.

How does Clef compare to Jev?

Per Cloudflare's own measurements Clef leads on 7 of 10 decision benchmarks, for example BANKING77 (94.20 versus 79.74 macro-F1) and BFCL (98.47 versus 95.75). On knowledge benchmarks Jev leads per the same table (GPQA Diamond 78.3 versus 48.0; MMLU-Pro 82.7 versus 65.9). An independent single-person test with 42 agent decisions saw Jev at 71.4 percent and Clef-flash at 66.7 percent with 83 percent agreement. Clef is compatible with the Jev API, so existing integrations can switch the base URL.

What is Cloudflare's RL fine-tuning platform?

Alongside Clef, Cloudflare announced a platform for reinforcement learning fine-tuning: AI Gateway captures data, Workers AI generates rollouts, Cloudflare Containers serve as the sandbox, a new Trainer updates the weights, and the result is redeployed via Workers AI with a bring-your-own model. At launch it is a guided offering with Cloudflare's forward-deployed engineers; a self-serve version is announced, pricing is not published.

Consultation for this model?

We help you select and integrate the right AI model for your use case.