@yibie: https://x.com/yibie/status/2104076142500069608

X AI KOLs Timeline Papers

Summary

This article demonstrates how to convert GLM-5.3-Flash into a Jev-style System 1 decision model, achieving typed decisions through a single forward pass. Benchmark tests show that it performs comparably to specialized models in terms of accuracy and speed.

https://t.co/GDNkb56KSn
Original Article
View Cached Full Text

Cached at: 09/27/26, 01:23 PM

Converting GLM-5.3-Flash into a Jev-style System One Model

Authors: Johannes Hötter and Dr. Marko Rosenmüller (Edgeless Systems / Privatemode)

Abstract: In this article, we demonstrate how an off-the-shelf LLM can make typed decisions with a single forward pass. This approach makes “turning an LLM into a Jev-style decision model” possible.

We evaluated this approach using GLM-5.3-Flash running on Privatemode. Using a benchmark constructed from public datasets, we show that this setup delivers results comparable to TypeSafe’s Jev in terms of decision accuracy and speed.

A bonus advantage: with this setup, GLM-5.3-Flash can make typed decisions on images, which Jev cannot do.

Why Typed Decisions Are Needed

Many things software asks an LLM are actually decisions. “Which team should handle this ticket?” or “Does this contract clause belong to the liability section?”

In these scenarios, software typically requires the LLM’s response to follow a specific format (like JSON) and to come from a predefined set (like “yes” and “no”).

With proper prompting, LLMs can often do this reliably already. However, with the basic approach, speed and cost become issues: for every decision, the LLM must write out an entire JSON object, and a reasoning model might have thought through hundreds of tokens before that. Additionally, you don’t get the model’s confidence (unless you explicitly ask for it). All these aspects can be critical in practice, and they have so far prevented people from using LLMs for decisions in high-concurrency scenarios.

Dedicated decision models like Jev and Laya (also called “System One” models) are designed to solve this problem. You pass in a state and a set of named options, and get back the selected option along with a confidence score (i.e., probability) for each option.

Turning an LLM into a Decision Model

We initially asked ourselves: can we turn an LLM into a decision model with properties like Jev’s? The short answer is: “Yes.”

Before understanding our approach, let’s understand how LLMs work:

LLMs never directly write text. Given a prompt, an LLM outputs a probability distribution over its entire vocabulary. In text generation, in the simplest case, the token with the highest probability is selected as the next token. The selected token is then appended to the prompt, and the process repeats. As mentioned, if you just want to set a few fields in a JSON object, this is slow and expensive.

Our core insight is: there’s no need to make the LLM predict the entire JSON object, because we already know its shape. We only care about the LLM’s typed judgment for a given input.

We realized we could construct a prompt that makes the LLM produce a typed judgment in one run—without fine-tuning, using the model as it comes out of the box. This is the difference from Jev and Laya—those are models specifically trained for this purpose. The basic steps are as follows:

  1. Number the options. The state, question, and output options are provided in the prompt as JSON, with each option assigned an index. The instruction asks the model to answer by prefixing with choice_index: followed by an index.
  2. Pre-fill the answer. The prompt ends with choice_index: . Thus, the first token generated by the model will be an index pointing to one of the predefined options.
  3. Evaluate the output. Instead of reading the token produced by the model, we read the probabilities it assigns to all option indices at that position. After normalizing over the options, these give the probability for each answer, and we simply take the most likely one.

The示意 is like this:

# End prompt within assistant response

user
{"state": "I was charged twice for my order.", "question": "Which team?", "options": [
  {"index": 0, "name": "payments"}
  {"index": 1, "name": "complaints"}
  {"index": 2, "name": "technical"}
]}
assistant
choice_index:

parameters:
max_tokens: 1
temperature: 0
allowed_token_ids

# Keep options, renormalize
0 payments   62.1%
1 complaints 37.7%
2 technical   0.2%
→ Answer: payments, confidence 0.39 (1 = certain, 0 = all options equally likely)

The prompt numbers the options and ends with the started assistant response choice_index: , so the next token is the index. A mask on the vocabulary restricts it to only the option indices, the API returns their log probabilities, and normalizing over the options gives the probability for each answer.

We implemented the above steps for GLM-5.3-Flash running on vLLM (inside Privatemode). We used the /chat/completions endpoint with continue_final_message and add_generation_prompt: false, as these parameters make the model continue from the pre-filled assistant turn in step 2, instead of starting a new one. They also allowed us to pass images alongside text—which is what enables typed decisions on images.

Three Practical Pitfalls

For vLLM and GLM-5.3-Flash, we found the following details important:

  1. allowed_token_ids can restrict the LLM’s output vocabulary to only the allowed options. It demotes every other token to -inf. We set it, but it acts as a guardrail, not a necessity.
  2. top_logprobs is insufficient for step 3. It reports the distribution before the restriction is applied, so format tokens like leading spaces occupy top slots, and some options drop out of the list, appearing to have zero probability. vLLM’s logprob_token_ids solves this: it returns log probabilities for the specific token IDs you specify.
  3. The token IDs for the indices depend on the model’s tokenizer. Digits are not always single tokens. For example, in GLM-5.3-Flash, 12 is a single token. Instead of shipping a model-specific tokenizer with the library, it fetches token IDs from the serving endpoint, making it easy to use with any model: sending a prompt with echo to /completions returns the exact tokenization the served model applies.

The implementation code is in the edgelesssys/privatemode-decisions repository.

Benchmark Results

We evaluated this approach using a self-built benchmark, also included in the repository.

We compared three systems on 29 public, labeled datasets: GLM-5.3-Flash querying on Privatemode using the above technique; TypeSafe’s Jev; and Convai’s Laya.

The datasets have between 2 and 151 options, covering intent routing, sentiment, topic classification, content moderation, entailment, question answering, legal text, and scanned documents. The corpora are in both English and German. All three systems received the same state, the same options in the same order, and the same instructions.

We ran each dataset twice. Even at temperature 0, our GLM-5.3-Flash and Jev changed up to 3.5% of answers between two identical runs: temperature 0 eliminates sampling randomness, but batch processing and floating-point arithmetic still make a single forward pass on a busy server not bitwise reproducible. Therefore, we treat smaller differences as noise.

We ran Jev and Laya with their default settings and did not tune our prompts on these datasets.

Accuracy

Comparing across 28 text datasets, GLM-5.3-Flash and Jev are neck-and-neck.

  • Each is more accurate on 10 datasets
  • On the remaining 8, the difference is within one percentage point
  • The median difference is 0.7 percentage points, slightly favoring Jev, which is not statistically significant (p = 0.64)
  • Laya (the 421M parameter model we ran locally) is lower than both on most datasets. Its median difference is 13 to 15 percentage points, which is statistically significant (p < 0.001)

The number of options affects accuracy more than the choice between Jev and GLM-5.3-Flash.

Three datasets labeled the same questions twice—once coarse, once fine—so the task remains the same, only the number of options changes:

TREC, from 6 to 42 options: Jev: 92.1% → 85.6% GLM-5.3-Flash: 91.2% → 79.6% Laya: 88.4% → 51.2%

MASSIVE, from 18 scenarios to 59 intents: Jev and GLM-5.3-Flash scores increase on both languages, while Laya’s scores decrease.

Thus, the number of options alone does not determine how difficult a task is.

Latency and Cost

We measured latency in separate runs, sending one request at a time—because timing under load measures the queue, not the model.

Since Privatemode is hosted in the EU and Jev in the US, we ran four datasets from both Germany and the US:

From Germany: Privatemode 180ms, Jev 264ms From the US: Jev 164ms, Privatemode 299ms

In terms of cost, Jev is cheaper. For one million decisions, GLM-5.3-Flash costs about €62, Jev about €16 (based on the respective service pricing).

An Interesting Finding About Tokens

One point per dataset: the number of input tokens GLM-5.3-Flash sends per question minus the tokens Jev sends for the same question.

Since the question text is the same for both, it cancels out; what remains is the difference in “wrapping.”

  • Jev adds a large fixed overhead, very little per option
  • GLM-5.3-Flash adds a small fixed overhead, more per option

The trend is: GLM-5.3-Flash starts about 219 tokens lower than Jev but adds about 11 tokens more per additional option. The breakeven point is around 21 options.

Multimodal Decisions

The state doesn’t have to be text. GLM-5.3-Flash is a vision-capable model, so a question can include an image—such as a scanned invoice, a photo of a damaged package, or a screenshot.

The image goes into the same prompt, and the answer is still a single token with probabilities for each option.

According to its documentation, Jev only works on text, and Laya is a text encoder, so neither accepts image input.

On RVL-CDIP (1,600 scanned business documents in 16 classes), GLM-5.3-Flash achieves 70.2% accuracy, and is the only one of the three that can answer. A document is more expensive than a sentence: an image adds about 1,350 input tokens, so one million document decisions cost about €270.

Capability Limits

As the number of options grows, each system hits limits.

Laya’s option names share a 192-token budget—enough for banking77’s 77 intents, but not for CLINC150’s 151 intents.

Privatemode’s deployment of GLM-5.3-Flash reports at most 128 items in logprob_token_ids, while the mask in allowed_token_ids can cover every option. So the library sends the problem with 151 options twice as is, reading probabilities for the first 128 options from the first response, and the other 23 from the second. Both requests run the same forward pass, so merging the results gives the distribution that a single request would return (within run-to-run noise).

On CLINC150, GLM-5.3-Flash achieves 87.5% this way, while Jev gets 78.4%. The second request and the long option list take time: one decision takes 719ms, compared to Jev’s 249ms.

Beyond that, the ceiling is the 191 option indices that GLM-5.3-Flash can fit into a single token.

Further Findings

  1. Some remaining errors are in the labels. When most systems agree on an answer and the dataset label disagrees, usually that label is one of two defensible answers. banking77 has many such pairs, like get_physical_card and order_physical_card, or declined_transfer and failed_transfer. About 17% of banking77 samples fall into this category, so the maximum any system can achieve there is around 85%, not 100%.
  2. Reasoning helps, but at a cost. As a control, we had the same GLM-5.3-Flash reason before answering on all 29 datasets. It is more accurate across each range of option counts—89.9% vs. 85.5% for 2 options, 82.0% vs. 79.2% for 21 to 80 options. But it writes hundreds of tokens per decision instead of one, costing about €350 per million decisions, compared to €62 for the single-token version. At the other extreme, pure embedding similarity (not using a decision model at all, just selecting the option whose text is closest) reaches 45.9% to 72.8%, depending on the range.
  3. Renaming options affects systems differently. We reran each dataset, replacing every option with a synonym, changing nothing else. On boolq, after renaming true and false to correct and wrong, GLM-5.3-Flash dropped by 20 points, while the other two dropped by less than three. (Author’s note: Because renaming also changes the meaning of the question, it does not isolate the factor of “memorization”.)

Build It Yourself

Our library is written in Python and can be used with any vLLM-backed endpoint; it depends on vLLM extensions to the OpenAI API, such as allowed_token_ids and logprob_token_ids.

The README explains how to configure Privatemode, and there’s an AGENTS.md file telling coding agents what an implementation must get right.

The benchmark repository contains the methodology, dataset specifications, the test framework, aggregation methods, and downloads of every raw run, so every number in this article can be reproduced without rerunning anything.

On Privatemode, these decisions are protected by confidential computing: your data remains encrypted in memory even during processing, and the client verifies the deployment’s attestation report before sending anything.

Links

Original: https://www.privatemode.ai/blog/system-one-from-glm-flash Implementation repository: https://github.com/edgelesssys/privatemode-decisions Benchmark repository (reproduce all numbers): https://github.com/edgelesssys/privatemode-decisions-benchmark Jev docs: https://docs.typesafe.ai/models Laya: https://huggingface.co/convaiinnovations/laya

#Jev #GLM #Benchmark

Similar Articles

Turning GLM-5.3-Flash into a Jev-like decision model

Hacker News Top

This article demonstrates how to adapt the GLM-5.3-Flash LLM to function as a typed decision model similar to Jev, achieving comparable accuracy and speed while enabling decisions on images in a single forward pass.

zai-org/GLM-5.3-Flash

Hugging Face Models Trending

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, with 320B total parameters and 18B active parameters, outperforming previous versions and approaching Claude Opus 4.8 through a redesigned hybrid architecture for improved efficiency.