I open sourced my hackathon search agent, but I’m still figuring out the best model for evaluation
Summary
The author open-sourced hackathon-searcher, a tool that discovers, evaluates, and helps apply to hackathons, and is seeking community advice on how to design its evaluation layer using open-source models.
Similar Articles
@seclink: Another open-source sample... Normally, you'd have to pay for this, but that said, since it's open-sourced, it's defini…
This post announces Autoresearch Bench, an open-source benchmark for coding agents to autonomously tackle research problems, noting stark differences between models in autoresearch loops.
Is it agentic enough? Benchmarking open models on your own tooling
This blog post introduces a benchmark methodology for evaluating how well open models perform on agentic coding tasks, focusing not just on accuracy but on the efficiency of the agent's process. It provides a customizable tooling harness using the pi coding agent and tests across models and library revisions.
I built a live ranking of every AI agent and foundation model (open source)
A developer launched AgentTape, a live ranking site that aggregates data from multiple sources (GitHub, Hugging Face, OpenRouter, etc.) to score and compare public AI agents and foundation models, aiming to provide a more holistic evaluation beyond benchmarks.
I built a cyberpunk agent matcher because I got tired of cloning broken repos
The author built a free quiz-based tool to help developers find suitable, well-maintained AI agent repositories among 50+ options, avoiding broken or outdated repos, and is seeking community feedback.
@_lamaahmad: We (@CedricWhitney, @SandhiniAgarwal, @EstherTetruas, @OliviaGWatkins2, @dgrobinson) wrote about nuances we’ve observed…
OpenAI researchers share lessons learned from working with third parties on frontier model evaluations, highlighting the importance of considering the evaluation harness and potential validity issues like reward hacking, contamination, and sandbagging.