Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time

Reddit r/LocalLLaMA Models

Summary

Mica v0.1 4B is an open-source decision model for AI agent loops, trained for under $30 on GPU time, designed to run locally on 8GB GPUs with good calibration and performance in specific tasks like prompt injection resistance.

I've been building a small decision model for agent loops: gates, routers, "should I ask the user or just act" checks. It's out now as Mica v0.1 4B (Apache-2.0). What it does You give it a state, a question and the allowed answers, and it returns a calibrated probability for each answer: yes/no, a choice among 2 to 255 options, or a score with 2 to 10 levels. It never generates text. It runs one prefill and reads the logits of the option labels at the answer position. It speaks the TypeSafe /v1/systemone format, so anything written for Jev works against it. How it's built - Qwen3.5-4B with a rank-16 LoRA on all 32 layers (attention and Gated DeltaNet), merged. No new heads, so it's a plain Qwen3.5-4B-shaped checkpoint. - About 34k source decisions, expanded to 77,732 training rows (about 34.7M tokens). Roughly half English and half Korean, across 12 areas: coding agents, code review, computer use, user requests, documents, policy rules, dates and quantities, routing, state tracking, games and general knowledge. - Plain cross-entropy on verified answers, one epoch, and one global temperature for calibration. - All experiments plus the final run cost under $30 of rented GPU time (RTX 3090s). Results Held-out set of 7,328 decisions, written after the training data was frozen and not opened until training finished. English subset, where every model can answer: - Jev 1.13 (closed API): 74.1 - Mica 4B: 67.0 - JevK5 4B: 61.0 - Kev 4B: 57.0 - Qwen3.5-4B base with the same readout: 55.0 Public sets, same prompt and readout for every model (Mica / Jev 1.13 / JevK5 / Kev 4B): - JevBench hard, public 111 items: 69.5 / 74.3 / 76.2 / 52.4 - SemIf: 94.4 / 98.4 / 86.1 / 89.3 - Kev transfer v9: 69.2 / 82.0 / 70.5 / 73.5 - MMLU-Pro, 10k items: 53.0 / 82.3 / 53.5 / 49.7 Through JevBench's official runner and the llama.cpp server, the public hard tier scores 64.9 instead of 69.5. I've submitted it for their sealed run. Where it's actually useful - In-data prompt injection. Put a note inside the state telling the judge to pick a wrong option, and Mica still gets 69% right (81% without the note). Jev drops to 18% and Kev to 31%. - Calibration. When it says 0.9 or higher, it's wrong 2.5% of the time on the held-out set (ECE 5.4%). - Local and small. The Q5_K_M file is 3.5 GB with no measurable accuracy loss against BF16 on our calibration set. Speed (RTX 3090, one request at a time, median over the 231 public JevBench items) - Mica Q4_K_M: 47 ms - Mica BF16: 54 ms - Kev 4B: 76 ms - JevK5 4B: 99 ms - Nimble 9B: 132 ms To be fair about this: the three 4B models share the same architecture, so most of the gap comes from the serving path, not the model. Mica ships as GGUF and runs on llama.cpp with a direct logits readout, while the others were measured through their own PyTorch code. On long inputs (around 3.7k tokens) Mica is slightly slower than JevK5. Limitations - Knowledge-heavy questions: MMLU-Pro 53 vs 82 for Jev. It's a 4B judge, not an encyclopedia. - Long English policy documents are its weakest public set. - Notes inside the state still nudge it. A note pointing at the right answer lifts accuracy to 89%. - It doesn't yet tell reversible from irreversible actions well. "Delete these files" and "move these files to trash" both get about 0.8 on "confirm first". - On harder reasoning items it's right but less sure than Jev (for example 0.55 vs 0.96 on a small ordering puzzle), so set your confidence thresholds accordingly. Try it Weights (BF16 safetensors and GGUF from Q4_0 to Q8_0): https://huggingface.co/sky7350/Mica-v0.1-4B Code, TypeSafe-compatible server and Docker setup: https://github.com/akivet/Mica-v0.1-4B The README has a one-line Docker command and a curl example. Happy to hear where it breaks. Ambiguous "act or ask" cases are what I most want to improve next.
Original Article

Similar Articles