Apodex releases open-weight small models (0.8B, 2B, 4B) specialized for agentic verification tasks, along with the AgentHarness evaluation framework for local agent workflows.
Hey r/LocalLLaMA, We just released **Apodex 1.0**, and alongside our flagship API, we are releasing the weights for our **Smol models (0.8B, 2B, and 4B)**. Our core research focuses on **independent verification** in long-horizon tasks. Instead of just scaling up parameter sizes for raw generation, we’ve been experimenting with small, highly specialized local models that handle specific sub-tasks in an agentic loop (like source cross-examination, hypothesis testing, and tool-grounded synthesis). We wanted to share the open weights and our evaluation harness with the community to get your thoughts on local agent workflows. # 🧠 The Setup: What are these Smol models for? When running long-horizon agents locally, using a massive 70B+ model for every single step (like checking if a URL is broken or verifying a regex) is incredibly inefficient. We specialized these 0.8B, 2B, and 4B models to act as sub-agents within our **AgentOS** runtime. They are trained to: 1. **Fact-check/Cross-examine:** Treat external text outputs as "claims" rather than ground truth. 2. **Execute & Verify:** Formulate precise tool calls and verify structural outputs before passing them back to the main controller. # 📊 Flagship Model Benchmarks (For Context) To give you an idea of what the full architecture is capable of when these verification loops are running at scale, our flagship model (**Apodex-1.0-H**) achieved the following scores: * **DeepSearchQA:** 94.4 | **BrowseComp:** 90.3 * **HLE-Text:** 60.8 * **SuperChem:** 74.2 * **FrontierScience Research:** 46.7 ( Frontier science reasoning is still a brutal bottleneck for all of us) # 🛠️ Open-Source Components & Local Evals We’ve open-sourced **AgentHarness**, which is the framework we use to test and evaluate these agentic workflows locally without drifting over 50+ steps. The open-weight models are hosted on Hugging Face, and the evaluation code is on GitHub. *(Note: To keep this post strictly compliant with the sub's rules, I’ve put all the Hugging Face links, GitHub repos, and the free early-access web platform in the stickied comment below).* **For those into local agent orchestration:** * Have you tried routing smaller tasks to <4B models in your local agent workflows? How do you mitigate the formatting/JSON adherence drift? * What are your thoughts on optimizing small models specifically for *verification* rather than conversational fluency? Would love to hear your feedback, and let me know if you want us to cook up some GGUF/EXL2 quants for these!
Apodex 1.0 is a self-evolving AI system post-trained on Qwen3.5, achieving SOTA on BrowseComp, DeepSearchQA, and HLE-text. Its 4B mini model outperforms 30B-class models, with an AgentOS runtime for task orchestration. Open weights available.
AgentOS and Apodex 1.0 introduce a runtime and open-weight model family for long-horizon agent tasks, using independent verification to prevent agent drift. The platform includes skeptical sub-agents and achieves high scores on complex benchmarks.
Apodex 1.1 improves sustained, verifiable progress on complex real-world tasks by scaling executable environments and training agents for long-horizon coordination, achieving leading performance with a smaller 35B-parameter model.
Apodex-1.0-H is a new deep research model that introduces a multi-agent architecture where the model decomposes tasks, spawns specialist sub-agents, and uses self-verification and iterative improvement to produce answers. Open-weight variants are available on HuggingFace.
Apodex is an open-sourced deep research harness and AI model, which is a finetune of Qwen 3.5 35B A3B, achieving performance comparable to frontier models with only 3B active parameters.