@MiaAI_lab: Cloudflare's Clef beat Jev in just two weeks How: "frozen" Qwen does one prefill pass, a tiny schema head scores every …

X AI KOLs Timeline Models

Summary

Cloudflare released Clef, an open-source 27B multimodal decision model that uses a frozen Qwen backbone and a tiny schema head to score all answer options in a single forward pass with no text generation, achieving 4x the speed and up to 2x the accuracy of Jev. A smaller faster variant, Clef-Flash, is also available on Hugging Face.

Cloudflare's Clef beat Jev in just two weeks How: "frozen" Qwen does one prefill pass, a tiny schema head scores every answer option in parallel, with no text generated at all. That's why it's 4x faster than Jev at 2x the accuracy on some benchmarks. And it's open source! Link to HF: https://huggingface.co/Cloudflare/clef
Original Article
View Cached Full Text

Cached at: 10/03/26, 04:53 AM

Cloudflare’s Clef beat Jev in just two weeks

How: “frozen” Qwen does one prefill pass, a tiny schema head scores every answer option in parallel, with no text generated at all.

That’s why it’s 4x faster than Jev at 2x the accuracy on some benchmarks.

And it’s open source!

Link to HF: https://huggingface.co/Cloudflare/clef


Cloudflare/clef · Hugging Face

Source: https://huggingface.co/Cloudflare/clef

Clef is a 27B multimodal model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question in a single forward pass. There is no free-form text generation and no output parsing.

The Clef API is fully compatible with Jev and SystemOne.

Clef is post-trained fromQwen/Qwen3.8-27B. SeeClef-Flashfor the smaller, faster variant.

https://huggingface.co/Cloudflare/clef#modelModel

  • **Backbone:**Qwen/Qwen3.8-27B with its vision encoder, stored as standard sharded safetensors.
  • **Joint schema head:**a small transformer head that reads the backbone’s final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly.
  • **Output:**one logit per allowed option for each question. Apply a softmax per question to get probabilities.

https://huggingface.co/Cloudflare/clef#filesFiles

FilePurposemodel\-\*\.safetensors,model\.safetensors\.index\.json,config\.json,generation\_config\.jsonBackbone, including the vision encoderjoint\_head\.safetensors,joint\_head\_config\.jsonJoint schema headjoint\_schema\_model\.pyRecord encoding, batching, the model,load\_release\_model, andsystemone``tokenizer\.json,tokenizer\_config\.json,chat\_template\.jinja,processor\_config\.jsonTokenizer and image/video processorLICENSEApache-2.0 license

https://huggingface.co/Cloudflare/clef#usageUsage

Tested withtorch2.11 andtransformers5.10.2 on a single H200. Image and video inputs also needpillow.

import sys

import torch
from huggingface_hub import snapshot_download

path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))

https://huggingface.co/Cloudflare/clef#jev–systemone-apiJev / SystemOne API

systemonetakes a Jev/SystemOnePOST /v1/systemonerequest body and returns the same response body:model,answerskeyed by question ID, andusage. Achoiceanswer haschoice,confidence, andprobabilities; ascoreanswer has the expectedscore,confidence,legend, andprobabilities; anoulanswer has the probability of true.instructionsis optional, andimagesandvideosmay be added to the request.

from joint_schema_model import systemone

response = systemone(model, processor, {
    "model": "clef",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

https://huggingface.co/Cloudflare/clef#images-and-videoImages and video

Addimages(PIL images) orvideos(frame arrays) to the record and pass the processor toencode\_record. Optional processor arguments go inmedia\_kwargs.

from PIL import Image

record = {
    "state": {"task": "Review the attached receipt."},
    "images": [Image.open("receipt.jpg")],
    "questions": {
        "legible": {"type": "noul", "instructions": "Is the receipt total legible?"},
    },
}
encoded = encode_record(processor.tokenizer, record, processor=processor)

Text-only and multimodal records can be mixed in the same batch.

https://huggingface.co/Cloudflare/clef#input-formatInput format

FieldDescriptionstateAny string or JSON value describing the situation to decide onimages,videosOptional lists of images or video frame arraysmedia\_kwargsOptional keyword arguments for the image/video processorquestionsMapping of question ID to question Each question has:

  • type:noul(true/false),choice(named options), orscore(ordered options)
  • instructions: what to decide; optional, and the question ID is used when it is omitted
  • criteria: forchoice, a mapping of option ID to description; forscore, a list of option descriptions indexed from 0; fornoul, optional descriptions fortrueandfalse

encode\_recordacceptsmax\_length(default 16,384 tokens) andmax\_state\_tokensto bound the input.

https://huggingface.co/Cloudflare/clef#resultsResults

https://huggingface.co/Cloudflare/clef#decision-indexDecision Index

Per-benchmark results from our internal run of theDecision Index0.2.1 suite. Scores are percentages; ForecastBench is a Brier score, where lower is better. The last two rows are request latency in milliseconds, where lower is better. The best value in each row is in bold.

BenchmarkClefClef-flashJevDiffusionGemma JevKev 9BLayaBFCL (case exact accuracy)98.598.895.896.594.538.1ToolRet (nDCG@10)69.266.465.361.264.312.8API-Bank (accuracy)91.993.188.283.756.311.5BANKING77 (macro-F1)94.290.979.774.384.814.3CLINC150+OOS (macro-F1)97.466.889.383.579.03.2RouterBench (selected quality)79.779.979.979.080.057.1Home appliance simulator (case exact accuracy)83.097.752.342.025.00.0SGD/SGD-X (macro-F1)43.834.243.040.664.042.4ContractNLI (macro-F1)81.484.371.776.057.829.0ANLI (macro-F1)69.859.174.866.456.348.7BPoMP (accuracy)96.995.490.686.967.051.6Humicroedit (accuracy)66.775.161.963.055.847.2POP909-CL (accuracy)15.81.618.12.510.85.1cfcolor (accuracy)66.065.864.758.256.352.3MMLU (accuracy)90.391.891.779.375.330.7GPQA Diamond (accuracy)48.051.078.344.938.827.6ARC-Easy (accuracy)99.099.599.398.297.747.0ARC-Challenge (accuracy)97.798.397.894.593.728.6WinoGrande (accuracy)93.597.592.073.673.250.5HellaSwag (accuracy)98.298.694.583.381.933.1GSM8K (accuracy)80.867.379.950.348.721.6ChessBench (accuracy)24.723.017.214.211.27.7MuSR (accuracy)83.586.066.161.257.943.2SATA-Bench (case exact accuracy)33.836.726.427.526.70.3BRIGHT (nDCG@10)45.939.347.542.938.519.9Amazon ESCI (macro-F1)57.557.455.253.449.224.4ACOS (per-review F1)33.325.929.524.518.33.5FinEntity (macro-F1)96.297.187.089.088.461.0VAST (macro-F1)59.549.664.655.755.440.5NLI4CT (macro-F1)82.978.684.178.474.947.7CRUXEval (accuracy)86.786.173.064.751.240.2CLadder (accuracy)94.097.772.667.862.052.9ForecastBench (Brier, lower is better)13.910.617.429.617.641.1Habermas Machine (accuracy)68.771.845.945.039.433.4PhishNChips (accuracy)79.675.062.585.450.750.1MMLU-Pro (accuracy)65.965.382.756.951.113.6BBH (accuracy)73.768.992.970.765.234.1RAGTruth (hallucination F1)79.435.676.570.446.248.8HoVer (accuracy)65.261.272.970.958.855.8When2Call MCQ (accuracy)72.465.681.075.449.611.9New Yorker (accuracy)69.566.170.163.658.127.1Median latency (ms)209.338.8524.184.451.45.8p95 latency (ms)238.6122.4536.0211.2187.9222.5

https://huggingface.co/Cloudflare/clef#workflow-evalsWorkflow evals

Decision accuracy on four end-to-end business workflows fromTypesafe Evals, scored against consensus reference labels. All models are scored on the same dataset revision and case cohort.

WorkflowMetricClefClef-flashJevInvoice processingExact actions64.757.161.8Invoice processingPrimary action86.273.383.1Customer serviceExact actions76.377.076.0Security incidentsExact actions62.961.761.7Agent trace observabilityPrimary action68.569.871.6

https://huggingface.co/Cloudflare/clef#licenseLicense

Released under the Apache-2.0 license, following the base modelQwen/Qwen3.8-27B.

Similar Articles

@victormustar: Alert: Cloudflare just dropped a Jev alternative on Hugging Face (Apache 2.0)

X AI KOLs Timeline

Cloudflare released Clef, an open-source (Apache 2.0) 27B multimodal decision model on Hugging Face that converts a state and a schema of typed questions into probabilities for all options in a single forward pass. Post-trained from Qwen3.8-27B, Clef has no free-form text generation and is API-compatible with Jev and SystemOne, with a smaller Clef-Flash variant available.

Cloudflare/clef-flash

Hugging Face Models Trending

Cloudflare released Clef-Flash, a 9B multimodal decision model built on Qwen3.5-9B that converts a state plus a schema of typed questions into per-option probabilities in a single forward pass, with no free-form text generation or output parsing.

Clef: Open Weights decision model by Cloudflare

Reddit r/LocalLLaMA

Cloudflare released Clef, an open-weights 27B multimodal decision model that takes a state and a schema of typed questions as input and returns probabilities for each option in a single forward pass, with no free-form generation or output parsing. It is post-trained from Qwen3.8-27B, ships on Hugging Face with a smaller Clef-Flash variant, and is compatible with the Jev/SystemOne API.

Clef

Product Hunt

Cloudflare introduces Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows, alongside a new reinforcement learning platform that lets developers fine-tune decision models with their own data.

Clef: our open-source decision models

Hacker News Top

Cloudflare 发布了两个开源决策模型 Clef 和 Clef-flash,托管于 Workers AI,采用 Apache 2.0 许可并在 Hugging Face 上开放下载,主打低成本、快速且一致的结构化输出,目前在 Jev Decision Index 榜单领先;同时推出新的强化学习微调平台,支持客户针对自身用例对 Clef 进行微调。