I open sourced my hackathon search agent, but I’m still figuring out the best model for evaluation

Reddit r/AI_Agents Tools

Summary

The author open-sourced hackathon-searcher, a tool that discovers, evaluates, and helps apply to hackathons, and is seeking community advice on how to design its evaluation layer using open-source models.

I've been going to a lot of hackathons recently and got tired of manually finding the good ones, researching travel support, checking deadlines and filling out similar applications over and over again. So I built hackathon-searcher, an open-source tool that can discover hackathons, research them, score them based on your preferences, help generate application answers and handle the application flow. The part I'm most interested in improving now is the evaluation layer. I want the system to take structured information about a hackathon, things like location, travel support, prizes, themes, eligibility and deadlines, combine that with a user's preferences and then decide how worthwhile that hackathon actually is for that person. Right now I'm trying to figure out the best architecture for this. Ideally I'd like to use an open-source model rather than relying entirely on paid APIs, but I'm not sure whether the best approach is: • one stronger open-source LLM doing the full evaluation • a smaller local model combined with deterministic scoring • having the model score individual criteria and calculating the final score separately • or using multiple models/evaluators and comparing the outputs I'm also trying to figure out which open-source models are actually good enough for this kind of structured judgment without making the system unnecessarily slow or expensive to run. The tool is still pretty early, so I'd especially love to hear from people building agents or LLM evaluation systems: how would you design this evaluation layer, and which open-source model would you use?
Original Article

Similar Articles

Is it agentic enough? Benchmarking open models on your own tooling

Hugging Face Blog

This blog post introduces a benchmark methodology for evaluating how well open models perform on agentic coding tasks, focusing not just on accuracy but on the efficiency of the agent's process. It provides a customizable tooling harness using the pi coding agent and tests across models and library revisions.