BeeNara is a 332MB ONNX cross-encoder model for document sorting that uses split-conformal prediction to provide calibrated 'none fits' responses for uncertain classifications.
Hey r/LocalLLaMA! I wanted to share a small project I’ve been working on called BeeNara Why I built this: I was looking for a way to automatically sort my local documents (invoices, letters, contracts) into my personal folders. While local LLMs are amazing, I noticed that smaller models (like Qwen3.5-4B) really struggle with one specific thing: admitting when a document doesn't fit into any of the provided categories. Instead of saying "I don't know", they tend to hallucinate and just shove the document into a random folder. Running a massive model just for basic sorting felt like overkill, especially on a laptop without a heavy GPU. What it does: BeeNara is a tiny (332 MB) ONNX cross-encoder model. You give it a document and a custom list of your folder names (like "Tax 2025" or "Invoices"), and it puts the document in the right one. The best part? It uses split-conformal prediction, meaning its confidence is highly calibrated. If it's not absolutely sure, or if none of your folders are a good match, it simply returns "none fits" and flags the document for human review. Key Features: Zero-shot: You just use plain text folder names. No fine-tuning or retraining needed. Fast & Local: Runs entirely offline on a laptop CPU in about 0.2–0.3 seconds per document (no PyTorch/GPU required, just ONNX runtime). Bilingual: Works seamlessly with English and German documents/folder names. High "None fits" recall: In benchmarks, it successfully catches 96.8% of documents where the correct folder is missing from the list. I originally built this as the category decider for a local document archivist tool, but you can easily use it standalone in Python. You can check out the model, code, and benchmark comparisons here: https://huggingface.co/Kwokou/BeeNara I'd love to hear your thoughts, feedback, or if you have ideas on how to improve it! Just wanted to share it with the community in case anyone else needs a fast, local "folder decider" that doesn't confidently lie to you. (Disclosure: I am the creator of this model!)
Introducing Ternary Bonsai 2 27B, a highly compressed AI model that retains 98.2% of performance while being 9x smaller in footprint, enabling efficient local deployment for tasks like reasoning, coding, and multimodal processing.
Bonsai 2 27B is a compressed AI model that achieves near-lossless performance in a 9x smaller footprint, with setup instructions provided for use with Prism's llama.cpp fork.
Ternary-Bonsai-2-27B is a 27B parameter AI model compressed into 2-bit ternary format, running on llama.cpp with CUDA and Metal support, facilitating lightweight on-device AI deployment.
Nanbeige4.2-3B is a compact agentic model using Looped Transformer architecture, demonstrating strong performance on agentic and reasoning benchmarks at the 3B scale, outperforming larger models such as Qwen3.5-9B and Gemma4-12B.
A tweet highlights Jina AI's ReaderLM-v2, a small 4GB model that achieves high accuracy in extracting information from messy DOM elements, exemplifying the trend toward specialized small language models.