Clef: Open Weights decision model by Cloudflare

Reddit r/LocalLLaMA Models

Summary

Cloudflare released Clef, an open-weights 27B multimodal decision model that takes a state and a schema of typed questions as input and returns probabilities for each option in a single forward pass, with no free-form generation or output parsing. It is post-trained from Qwen3.8-27B, ships on Hugging Face with a smaller Clef-Flash variant, and is compatible with the Jev/SystemOne API.

No content available
Original Article
View Cached Full Text

Cached at: 10/01/26, 06:39 PM

Cloudflare/clef · Hugging Face

Source: https://huggingface.co/Cloudflare/clef

Clef is a 27B multimodal model that turns a state and a schema of typed questions into decisions. It reads the state as text, JSON, images, or video, and returns a probability for every allowed option of every question in a single forward pass. There is no free-form text generation and no output parsing.

The Clef API is fully compatible with Jev and SystemOne.

Clef is post-trained fromQwen/Qwen3.8-27B. SeeClef-Flashfor the smaller, faster variant.

https://huggingface.co/Cloudflare/clef#modelModel

  • **Backbone:**Qwen/Qwen3.8-27B with its vision encoder, stored as standard sharded safetensors.
  • **Joint schema head:**a small transformer head that reads the backbone’s final hidden states, routes evidence from the state to each question, and scores all options of all questions jointly.
  • **Output:**one logit per allowed option for each question. Apply a softmax per question to get probabilities.

https://huggingface.co/Cloudflare/clef#filesFiles

FilePurposemodel\-\*\.safetensors,model\.safetensors\.index\.json,config\.json,generation\_config\.jsonBackbone, including the vision encoderjoint\_head\.safetensors,joint\_head\_config\.jsonJoint schema headjoint\_schema\_model\.pyRecord encoding, batching, the model,load\_release\_model, andsystemone``tokenizer\.json,tokenizer\_config\.json,chat\_template\.jinja,processor\_config\.jsonTokenizer and image/video processorLICENSEApache-2.0 license

https://huggingface.co/Cloudflare/clef#usageUsage

Tested withtorch2.11 andtransformers5.10.2 on a single H200. Image and video inputs also needpillow.

import sys

import torch
from huggingface_hub import snapshot_download

path = snapshot_download("Cloudflare/clef")
sys.path.insert(0, path)
from joint_schema_model import collate_records, encode_record, load_release_model

model, processor = load_release_model(path, device="cuda")

record = {
    "state": {"invoice": {"vendor": "Acme", "total": 1250.0, "currency": "USD", "status": "overdue"}},
    "questions": {
        "status": {
            "type": "choice",
            "instructions": "What is the invoice status?",
            "criteria": {"paid": "Invoice is paid.", "overdue": "Invoice is past due.", "draft": "Not sent."},
        },
        "large": {"type": "noul", "instructions": "Is the total above 1000 USD?"},
    },
}

encoded = encode_record(processor.tokenizer, record, processor=processor)
batch = collate_records([encoded], processor.tokenizer.pad_token_id, torch.device("cuda"))
with torch.inference_mode():
    logits = model(batch)[0]

for question, question_logits in zip(encoded.questions, logits):
    probabilities = question_logits.float().softmax(-1).tolist()
    print(question.question_id, dict(zip(question.option_ids, probabilities)))

https://huggingface.co/Cloudflare/clef#jev–systemone-apiJev / SystemOne API

systemonetakes a Jev/SystemOnePOST /v1/systemonerequest body and returns the same response body:model,answerskeyed by question ID, andusage. Achoiceanswer haschoice,confidence, andprobabilities; ascoreanswer has the expectedscore,confidence,legend, andprobabilities; anoulanswer has the probability of true.instructionsis optional, andimagesandvideosmay be added to the request.

from joint_schema_model import systemone

response = systemone(model, processor, {
    "model": "clef",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
        "outage": {"type": "noul", "instructions": "Is a service down?"},
    },
})
print(response["answers"])

https://huggingface.co/Cloudflare/clef#images-and-videoImages and video

Addimages(PIL images) orvideos(frame arrays) to the record and pass the processor toencode\_record. Optional processor arguments go inmedia\_kwargs.

from PIL import Image

record = {
    "state": {"task": "Review the attached receipt."},
    "images": [Image.open("receipt.jpg")],
    "questions": {
        "legible": {"type": "noul", "instructions": "Is the receipt total legible?"},
    },
}
encoded = encode_record(processor.tokenizer, record, processor=processor)

Text-only and multimodal records can be mixed in the same batch.

https://huggingface.co/Cloudflare/clef#input-formatInput format

FieldDescriptionstateAny string or JSON value describing the situation to decide onimages,videosOptional lists of images or video frame arraysmedia\_kwargsOptional keyword arguments for the image/video processorquestionsMapping of question ID to question Each question has:

  • type:noul(true/false),choice(named options), orscore(ordered options)
  • instructions: what to decide; optional, and the question ID is used when it is omitted
  • criteria: forchoice, a mapping of option ID to description; forscore, a list of option descriptions indexed from 0; fornoul, optional descriptions fortrueandfalse

encode\_recordacceptsmax\_length(default 16,384 tokens) andmax\_state\_tokensto bound the input.

https://huggingface.co/Cloudflare/clef#resultsResults

https://huggingface.co/Cloudflare/clef#decision-indexDecision Index

Per-benchmark results from our internal run of theDecision Index0.2.1 suite. Scores are percentages; ForecastBench is a Brier score, where lower is better. The last two rows are request latency in milliseconds, where lower is better. The best value in each row is in bold.

BenchmarkClefClef-flashJevDiffusionGemma JevKev 9BLayaBFCL (case exact accuracy)98.598.895.896.594.538.1ToolRet (nDCG@10)69.266.465.361.264.312.8API-Bank (accuracy)91.993.188.283.756.311.5BANKING77 (macro-F1)94.290.979.774.384.814.3CLINC150+OOS (macro-F1)97.466.889.383.579.03.2RouterBench (selected quality)79.779.979.979.080.057.1Home appliance simulator (case exact accuracy)83.097.752.342.025.00.0SGD/SGD-X (macro-F1)43.834.243.040.664.042.4ContractNLI (macro-F1)81.484.371.776.057.829.0ANLI (macro-F1)69.859.174.866.456.348.7BPoMP (accuracy)96.995.490.686.967.051.6Humicroedit (accuracy)66.775.161.963.055.847.2POP909-CL (accuracy)15.81.618.12.510.85.1cfcolor (accuracy)66.065.864.758.256.352.3MMLU (accuracy)90.391.891.779.375.330.7GPQA Diamond (accuracy)48.051.078.344.938.827.6ARC-Easy (accuracy)99.099.599.398.297.747.0ARC-Challenge (accuracy)97.798.397.894.593.728.6WinoGrande (accuracy)93.597.592.073.673.250.5HellaSwag (accuracy)98.298.694.583.381.933.1GSM8K (accuracy)80.867.379.950.348.721.6ChessBench (accuracy)24.723.017.214.211.27.7MuSR (accuracy)83.586.066.161.257.943.2SATA-Bench (case exact accuracy)33.836.726.427.526.70.3BRIGHT (nDCG@10)45.939.347.542.938.519.9Amazon ESCI (macro-F1)57.557.455.253.449.224.4ACOS (per-review F1)33.325.929.524.518.33.5FinEntity (macro-F1)96.297.187.089.088.461.0VAST (macro-F1)59.549.664.655.755.440.5NLI4CT (macro-F1)82.978.684.178.474.947.7CRUXEval (accuracy)86.786.173.064.751.240.2CLadder (accuracy)94.097.772.667.862.052.9ForecastBench (Brier, lower is better)13.910.617.429.617.641.1Habermas Machine (accuracy)68.771.845.945.039.433.4PhishNChips (accuracy)79.675.062.585.450.750.1MMLU-Pro (accuracy)65.965.382.756.951.113.6BBH (accuracy)73.768.992.970.765.234.1RAGTruth (hallucination F1)79.435.676.570.446.248.8HoVer (accuracy)65.261.272.970.958.855.8When2Call MCQ (accuracy)72.465.681.075.449.611.9New Yorker (accuracy)69.566.170.163.658.127.1Median latency (ms)209.338.8524.184.451.45.8p95 latency (ms)238.6122.4536.0211.2187.9222.5

https://huggingface.co/Cloudflare/clef#workflow-evalsWorkflow evals

Decision accuracy on four end-to-end business workflows fromTypesafe Evals, scored against consensus reference labels. All models are scored on the same dataset revision and case cohort.

WorkflowMetricClefClef-flashJevInvoice processingExact actions64.757.161.8Invoice processingPrimary action86.273.383.1Customer serviceExact actions76.377.076.0Security incidentsExact actions62.961.761.7Agent trace observabilityPrimary action68.569.871.6

https://huggingface.co/Cloudflare/clef#licenseLicense

Released under the Apache-2.0 license, following the base modelQwen/Qwen3.8-27B.

Similar Articles

Cloudflare/clef-flash

Hugging Face Models Trending

Cloudflare released Clef-Flash, a 9B multimodal decision model built on Qwen3.5-9B that converts a state plus a schema of typed questions into per-option probabilities in a single forward pass, with no free-form text generation or output parsing.

@victormustar: Alert: Cloudflare just dropped a Jev alternative on Hugging Face (Apache 2.0)

X AI KOLs Timeline

Cloudflare released Clef, an open-source (Apache 2.0) 27B multimodal decision model on Hugging Face that converts a state and a schema of typed questions into probabilities for all options in a single forward pass. Post-trained from Qwen3.8-27B, Clef has no free-form text generation and is API-compatible with Jev and SystemOne, with a smaller Clef-Flash variant available.

Clef

Product Hunt

Cloudflare introduces Clef and Clef-flash, open-source decision models hosted on Workers AI for high-speed classification and agentic workflows, alongside a new reinforcement learning platform that lets developers fine-tune decision models with their own data.

Clef: our open-source decision models

Hacker News Top

Cloudflare 发布了两个开源决策模型 Clef 和 Clef-flash,托管于 Workers AI,采用 Apache 2.0 许可并在 Hugging Face 上开放下载,主打低成本、快速且一致的结构化输出,目前在 Jev Decision Index 榜单领先;同时推出新的强化学习微调平台,支持客户针对自身用例对 Clef 进行微调。