Cloudflare Clef: Open Jev Alternative for Text, Image & Video

Cloudflare Clef and Clef-Flash are open-weight decision models that read text, images and video. Benchmarks, Workers AI pricing, how to run them, and how they compare with Jev and Laya.

Last updated: Oct 2, 2026

Cloudflare Clef is an open-weight decision model from Cloudflare. You send it a state (text, JSON, images or video) and a set of typed questions, and Clef returns a probability for every allowed option of every question in one forward pass, with no generated text to parse. Cloudflare released two sizes on October 1, 2026: Clef (27B) and the faster Clef-Flash (9B), both under Apache 2.0 on Hugging Face and both hosted on Workers AI.

Clef belongs to the same category as TypeSafe's Jev and the open-source Laya: System One models, which make fast, bounded decisions while a larger model does the slow reasoning. Its request and response format follows Jev's, and unlike Jev or Laya it reads images and video as well as text. This page collects what Cloudflare has published, how to run Clef, and how it compares with Jev, Laya and the other models released the same week.

Cloudflare Clef at a glance

ClefClef-Flash
ReleasedOctober 1, 2026October 1, 2026
Base modelQwen3.8-27BQwen3.5-9B
Parameters27B9B
Download sizeabout 55 GBabout 19 GB
LicenseApache 2.0Apache 2.0
Inputtext, JSON, images, videotext, JSON, images, video
Question typeschoice, score, noulchoice, score, noul
Workers AI price$0.24 per 1M input tokens$0.09 per 1M input tokens
Workers AI limits65,536-token context, 1 to 64 questions, up to 4 imagessame
Median latency209.3 ms38.8 ms

Sources: the Clef and Clef-Flash model cards, the Workers AI model page and the Cloudflare announcement. Latency is Cloudflare's own measurement over its benchmark suite.

How Clef works

Clef keeps the Qwen backbone, including its vision encoder, and adds a small joint schema head. The head reads the backbone's final hidden states, routes evidence from the state to each question, and scores all options of all questions together. The output is one logit per allowed option; a softmax per question turns it into probabilities.

That design is why Clef can answer many questions about the same input in a single pass: the state is read once, not once per question. According to the announcement, Cloudflare trained the head and low-rank adapters on a frozen backbone, using cross-entropy and Brier loss plus a reinforcement-learning step called RLCD (Reinforcement Learning for Calibrated Decisions, the same approach behind Jev and Laya), aimed at probabilities that match how often the model is right.

Cloudflare says it built Clef for its own decision-heavy work, such as Trust & Safety review, support triage and bot classification.

Benchmarks: what Cloudflare reported

Cloudflare ran its own copy of the Decision Index suite (more than 40 tasks) on Clef, Clef-Flash, Jev, Kev 9B, Laya and others. A selection:

BenchmarkClefClef-FlashJevLaya
BANKING77 (macro-F1)94.290.979.714.3
CLINC150 + out-of-scope (macro-F1)97.466.889.33.2
BFCL (exact accuracy)98.598.895.838.1
ContractNLI (macro-F1)81.484.371.729.0
RAGTruth (hallucination F1)79.435.676.548.8
MMLU (accuracy)90.391.891.730.7
GPQA Diamond (accuracy)48.051.078.327.6
BBH (accuracy)73.768.992.934.1
ForecastBench (Brier, lower is better)13.910.617.441.1
Median latency (ms)209.338.8524.15.8
p95 latency (ms)238.6122.4536.0222.5

On four end-to-end business workflows from Typesafe Evals, Clef beat Jev on invoice processing (86.2 against 83.1 for the primary action) and security incidents (62.9 against 61.7), was slightly ahead on customer service (76.3 against 76.0), and trailed on agent trace observability (68.5 against 71.6).

How to read these numbers:

  • They are self-reported. Cloudflare ran every model itself; no independent run has been published yet.
  • Jev still wins on hard reasoning. GPQA Diamond, BBH and MMLU-Pro all favour Jev by a wide margin.
  • Clef-Flash is not just a weaker Clef. It beats the 27B model on many tasks, but falls apart on some (CLINC150 with out-of-scope, RAGTruth), so test it on your own data.
  • Laya is far behind on accuracy and far ahead on speed. Laya is a 421M encoder: it has little world knowledge for tasks like MMLU or GPQA, and Cloudflare did not say which Laya checkpoint or settings it used. Its 5.8 ms median is the fastest in the table. For a test on everyday decision tasks with the same questions for every model, see our System One benchmark.

How to use Cloudflare Clef

On Workers AI

Clef is available on Workers AI as @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Inside a Worker:

const response = await env.AI.run('@cf/cloudflare/clef', {
  model: 'clef',
  state: 'Checkout has been failing for every customer for the last hour.',
  questions: {
    urgent: { type: 'noul', instructions: 'Is this support request urgent?' },
    team: {
      type: 'choice',
      instructions: 'Which team should handle this request?',
      criteria: {
        billing: 'Payments, invoices, and refunds',
        technical: 'Outages, errors, and configuration',
      },
    },
    severity: {
      type: 'score',
      instructions: 'How severe is the customer impact?',
      criteria: ['No impact', 'Minor', 'Major', 'Critical'],
    },
  },
});

From anywhere else, call the REST endpoint with a Cloudflare API token:

curl https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/ai/run/@cf/cloudflare/clef-flash \
  -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
  -d '{"model": "clef-flash", "state": "...", "questions": {"spam": {"type": "noul", "instructions": "Is this message spam?"}}}'

The request body is the same state + questions shape that Jev uses, plus an optional images list. The URL and the auth header are Cloudflare's, so code that builds the Jev request body itself only needs a new endpoint and token; check before relying on a Jev SDK.

On your own GPU

The weights and inference code are on Hugging Face. Cloudflare tested them with torch 2.11 and transformers 5.10.2 on a single H200. The systemone helper takes a Jev request body and returns a Jev response body:

import sys
from huggingface_hub import snapshot_download

path = snapshot_download('Cloudflare/clef-flash')
sys.path.insert(0, path)
from joint_schema_model import load_release_model, systemone

model, processor = load_release_model(path, device='cuda')
response = systemone(model, processor, {
    'model': 'clef-flash',
    'state': 'Our checkout started returning errors and orders are blocked.',
    'questions': {
        'outage': {'type': 'noul', 'instructions': 'Is a service down?'},
    },
})
print(response['answers'])

Hardware: Clef downloads about 55 GB of weights, so plan for a single 80 GB-class GPU such as an H100 or H200. Clef-Flash is about 19 GB, so it needs a GPU with at least 24 GB of memory (our estimate from the file size, before activations). Neither is meant for CPU. If you need decisions on a CPU, a laptop or in the browser, Laya runs in about 3 GB of RAM.

Clef vs Jev, Laya and the other new decision models

Clef arrived in a crowded week: Perplexity and AWS also released open decision models between September 30 and October 1, 2026.

Cloudflare Clef / Clef-FlashTypeSafe JevPerplexity pplx-decider-v1-27bAWS Strands Decider 2BLaya
Weightsopen, Apache 2.0closedopen, Apache 2.0open, Apache 2.0 (LoRA adapter)open, Apache 2.0
Size27B / 9Bundisclosed27B2B421M / 322M
Inputtext, JSON, images, videotexttext, imagestexttext, JSON
Hosted optionWorkers AI, $0.24 / $0.09 per 1M input tokensTypeSafe, $0.042 per 1M input tokensPerplexity APInonehosted Laya API or self-host
Self-host hardware80 GB / 24 GB GPUnot possibleGPU with about 49 GiB for weightssmall GPU (2B base)CPU, GPU or Apple Silicon
Jev request formatyesreferenceown Decider classnot statedyes, through laya-serve

All accuracy claims in this category are still self-reported and measured on different suites. See the full list in System One models.

Which one should you use?

  • You need images or video in the decision, such as screenshots, receipts or product photos: Clef or Clef-Flash. Jev and Laya read text only.
  • You already run on Cloudflare: Clef-Flash on Workers AI is the shortest path, with no GPU to manage.
  • You want Jev-level accuracy on open weights and own a large GPU: Clef, and compare it with pplx-decider on your own labelled data.
  • You need the lowest latency and cost, or no GPU at all: Laya. It answers in milliseconds on a CPU and can be fine-tuned on your own labels, which closes much of the accuracy gap on a narrow task.
  • You want the hardest reasoning questions right: Jev still leads on GPQA, BBH and MMLU-Pro in Cloudflare's own numbers.

FAQ

Is Cloudflare Clef open source?

Yes. Both Clef and Clef-Flash are released under Apache 2.0 on Hugging Face, with the weights, the joint schema head and the inference code.

Is Cloudflare Clef free?

The weights are free to download and run on your own hardware. On Workers AI you pay per input token: $0.24 per million for Clef and $0.09 per million for Clef-Flash.

Is Clef compatible with the Jev API?

The request and response bodies follow Jev's format: a state, typed choice, score and noul questions, and answers with probabilities. On Workers AI the endpoint URL and authentication are Cloudflare's, so you change the URL and token rather than only a model name.

What is the difference between Clef and Clef-Flash?

Clef is 27B and built on Qwen3.8-27B; Clef-Flash is 9B and built on Qwen3.5-9B. Clef-Flash is about five times faster (38.8 ms against 209.3 ms median in Cloudflare's tests) and cheaper on Workers AI, and it matches or beats Clef on many benchmarks, but it drops sharply on a few, such as out-of-scope intent detection.

Should I use Clef or Laya?

Use Clef when accuracy on varied tasks or image input matters and you have a large GPU or use Workers AI. Use Laya when you need decisions in milliseconds on ordinary hardware, want to keep data on your own machines without a GPU, or plan to fine-tune a small model on your own labels.