@yibie: https://x.com/yibie/status/2104024011692753228

X AI KOLs Timeline Tools

Summary

This article explains how to use the logprobs parameter in LLMs to implement Jev-style encapsulation for supporting structured queries, and extend to visual models via the attachments field for image processing, offering a flexible and customizable computer vision approach.

https://t.co/yrP26nAB9A
Original Article
View Cached Full Text

Cached at: 09/27/26, 11:20 AM

A Jev-style Wrapper for LLMs—Even Vision Models Can Use It

Author: Allan Riordan Boll (works at Google, blog allanrbo)

Jev and the self-hostable projects around it (like OpenJev and SemIf) intrigued me. Reading them taught me a clever trick: reading LLM token probabilities.

Obviously, this is an old trick for some (e.g., OpenAI’s logprobs cookbook), but it was new to me.

The Basic Idea

I think the core idea is to write a prompt like this:

State: My order arrived broken, I want a refund. Question: Which team should handle this? [A] billing [B] shipping [C] returns Answer with only the letter of the best option.

Then add a few JSON parameters in a compatible Chat Completions request:

{
  "max_completion_tokens": 1,
  "logprobs": true,
  "top_logprobs": 20
}

The LLM API will return that letter, along with the log probabilities of other candidate tokens the model considered.

Repeat this process for each question. Forcing it to generate only one token avoids verbose answers and is super fast—though processing the input still takes time. But if the backend supports it, each question can share the same KV cache for the state prefix.

The Fun Part: Vision Models Can Use It Too

Jev’s current documentation only describes a text/JSON state format. In my local experiments, I added an attachments field to include images.

My example: capture webcam frames, send base64 JPEG, then print a table—whether a person is present, indoor or outdoor, and how bright the scene is.

Running Gemma 4 12B on my RTX 3090, I get about 1 frame per second, asking three questions per frame.

I also ran it against OpenAI’s gpt-6-luna at about 0.2 FPS. That’s probably because I didn’t try to avoid the overhead of opening a separate connection to their system for each question per frame.

Dedicated computer vision models are certainly much more efficient, but I like the flexibility here: changing a judgment criteria only requires describing it in plain language.

The Code (Key Points)

The author provided a complete, runnable Python script. The key parts:

The request format (his extension):

{
  "state": "Inspect this webcam frame. Judge only what is visibly present.",
  "attachments": [],
  "questions": {
    "person": { "type": "noul", "instructions": "Is a person visible?" },
    "plant": { "type": "noul", "instructions": "Is a plant visible?" },
    "setting": {
      "type": "choice",
      "instructions": "Where is the camera?",
      "criteria": { "indoors": null, "outdoors": null, "unclear": null }
    },
    "light": {
      "type": "score",
      "instructions": "How bright is the scene?",
      "criteria": ["dark", "dim", "bright"]
    }
  }
}

attachments is his extension to the Jev request format—can be image paths or base64 data URLs. They are loaded once and shared across all questions.

The three question types map to letter options:

  • choice → directly uses the provided criteria
  • noul → {"true": ..., "false": ...}
  • score → numbers each level as "0", "1", "2", …

Then a key constraint unifies them:

Ask it for only one option letter, so the log probability of that letter represents that option.

The request body has two variants (OpenAI needs the Responses interface to get enough alternatives; llama.cpp uses the Chat interface):

# OpenAI
endpoint = "/responses"
body = {
    "model": model,
    "input": [{"role": "user", "content": content}],
    "reasoning": {"effort": "none"},
    "max_output_tokens": 16,
    "top_p": 1,
    "top_logprobs": 20,
    "include": ["message.output_text.logprobs"],
}

# llama.cpp / compatible interfaces
endpoint = "/chat/completions"
body = {
    "model": model,
    "messages": [{"role": "user", "content": content}],
    "max_completion_tokens": 1,
    "temperature": 0,
    "reasoning_effort": "none",
    "logprobs": True,
    "top_logprobs": 1024,
}

Note top_p: 1—this avoids cutting off alternatives.

Normalize the returned scores, with a robust detail handling:

peak = max(logprobs[letter] for letter in letters if letter not in missing)
weights = [math.exp(logprobs[letter] - peak) if letter not in missing else 0 for letter in letters]
total = sum(weights)

Then comes a defensive check worth learning:

An omitted token cannot rank before the last returned alternative. It’s only allowed to be zero if their combined normalized probability is below 1e-6.

If an option is missing, check:

missing_weight / (total + missing_weight) >= 1e-6

→ Then raise an error: “API omitted a non-negligible option score”

The purpose of this check is: if the API silently omits the score for some option, you end up with an incorrect probability distribution—and the error is silent. He chooses to raise an error directly, rather than pretending that option has zero probability.

Finally, restore the answer based on question type:

  • choice → the one with the highest probability
  • noul → probabilities["true"]
  • score → expected level (sum of level numbers multiplied by their probabilities)

The Main Loop

Camera preview and a background worker run in parallel—the worker scores one frame at a time:

executor = concurrent.futures.ThreadPoolExecutor(max_workers=1)

OpenCV is used only to capture the camera feed, not for any actual computer vision. The image goes through: camera → base64 → request body → logprobs → table.

My Assessment

This is one of the few posts in the Jev line that “thoroughly explains the mechanism and provides runnable code.”

It has three concrete values:

  1. It distills “the Jev thing” back to two JSON parameters: logprobs plus top_logprobs. Any Chat Completions-compatible endpoint can do this—you don’t need Jev’s API, nor even a specialized model.
  2. Vision is the step that suddenly broadens the applicability of this approach. Jev’s official docs only cover text/JSON state. Adding an attachments field allows the same scoring logic to handle images—and those judgments (person present? indoors/outdoors? how bright?) can be changed anytime in plain language, without retraining the model.
  3. The check for “API omitted option” is the most easily overlooked pitfall in this approach. Everyone will write “normalize then take the max,” but few will check “whether the denominator I received is complete.” His handling is worth copying directly.

Honest Boundaries: The author himself says dedicated CV models are much more efficient, and 1 FPS on a 3090 or 0.2 FPS on OpenAI aren’t production-ready numbers. Its value lies in flexibility, not performance.

Links

Original post: http://allanrbo.blogspot.com/2026/09/a-jev-like-wrapper-for-llms-including.html OpenAI’s logprobs cookbook: https://cookbook.openai.com/examples/using_logprobs OpenJev: https://github.com/TheoLeeCJ/SemIf-OpenJev

#Jev #logprobs #VisionModels

Similar Articles

@yibie: https://x.com/yibie/status/2104076142500069608

X AI KOLs Timeline

This article demonstrates how to convert GLM-5.3-Flash into a Jev-style System 1 decision model, achieving typed decisions through a single forward pass. Benchmark tests show that it performs comparably to specialized models in terms of accuracy and speed.

@yibie: https://x.com/yibie/status/2102914640707567798

X AI KOLs Timeline

This article tests the application of the Jev model in RAG retrieval, evaluating its effects on accelerating retrieval and reducing costs. Results show advantages in reranking and judging answerability.