ImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench

Reddit r/LocalLLaMA Models

Summary

A business consultant fine-tuned a 4B parameter AI model called ImaJev-4b for making decisions from text and images, achieving top performance on JevBench and DecisionBench benchmarks.

Some context first. I am process improvement / business consultant and had worked with Fortune 500 companies on improving their processes around refunds, returns, customer support, etc. This entire thing had a lot of complex decision making and generally every decision / condition node in a process map was generally replaced by a human because factoring in ambiguity in a code is very difficult. Idea of ImaJev Hence, when Jev came out, I was very intrigued with it and also could clearly see its use-case of improving decision making in complex decision work flows. However, Jev didnt have support for Images and I thought that it can be replicated for both Text and Images in a single model and thats when I started with ImaJev. Training Process It went badly at first. My first big fine-tune on about 500k short decisions made the 9B model worse at reasoning: 64.9 down to 42.3 on JevBench hard. It had basically learned to pattern-match. I spent the next couple of weeks generating hard questions with open-weight models and only keeping the ones where two Ai models agreed on the answer. That brought it back. Results Then, on the JevBench - It came out #1 of 91 (v1.4.2.2, scored 27 Sep), 67.37 vs Jev 1.13.0 at 63.29. The same week DecisionBench put it #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1. I honestly didn't expect either. To be fair about it: the #1 is on a score that weighs accuracy, calibration, speed and cost equally. On accuracy alone it's #3. Its main strength is that when it says 90% it's usually right, and it'll say "can't tell" instead of guessing. What it actually is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions. It gives back a probability for each option plus "unknown", in one forward pass. Runs on a Mac with MLX or on one GPU. The whole project costed me around $1200 in rented GPU and a lot of time :P Weights (Apache-2.0): https://huggingface.co/mohit67890/imajev-4b Code: https://github.com/mohit67890/imajev Demo: https://huggingface.co/spaces/mohit67890/imajev JevBench board: https://benchmarkheaven.com/jev-models DecisionBench board: https://huggingface.co/spaces/Hanno-Labs/decision-bench-leaderboard I would love to know your thoughts on it - it anyone would be interested to try that.
Original Article

Similar Articles