@xdotli: people came to our discord and ask how to write good proposal for making an eval when i started skillsbench, we have 0 …
Summary
SkillsBench founder shares the project's rapid growth from zero to 1600+ Discord members, 2 papers, and 150+ citations in under six months, along with extensive documentation.
View Cached Full Text
Cached at: 06/28/26, 10:06 AM
people came to our discord and ask how to write good proposal for making an eval
when i started skillsbench, we have 0 community member, 0 paper, 0 citation. now we have 1600+ members, 2 papers, 150+ citations in <6mos
sharing the whole SkillsBench docs + 200 pages of notes
https://docs.google.com/document/d/17f_qDeYPaNQRVDIFIr5topEUMd4_hv1RboVGGLGgdLc/edit?tab=t.6suyivndnk76#heading=h.9xcw13hxev1n…
thanks for your help @xdotli i m restructuring my application
Similar Articles
@Steve_Yegge: SkillBench is one of the most crazily important startups I know about, and it's been tough not to talk about them. Cong…
Steve Yegge praises SkillBench, a startup that scans coding agent session traces to build skills profiles and improve token efficiency. Matt Beane announces he's stepping in as full-time CEO.
@gneubig: SkillsBench is a great benchmark, and I'm not just saying that because @OpenHandsDev beats everyone else on it Skills a…
SkillsBench 1.1 is a benchmark for evaluating AI agents' ability to use skills, now fully audited and error-free.
@xdotli: A big pain point in using AI benchmarks is encountering errors after its first release. Today, we're releasing SkillsBe…
SkillsBench 1.1 is released as the first audited, error-free benchmark for AI agent skills, showing rapid capability improvement from ~36% to 67% resolution rate and demonstrating that skills can substitute for model scale.
@vincentsunnchen: New Benchtalks with @jyangballin: on ProgramBench (0% frontier models at launch) and the lineage/future of coding bench…
A podcast/interview episode discussing ProgramBench, a new coding benchmark where frontier models scored 0% at launch, covering its design philosophy, artifact-level evaluation, and the evolution of coding benchmarks from SWE-bench and InterCode.
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills
SkillEvolBench is a diagnostic benchmark for evaluating whether large language model agents can distill episodic experience into reusable procedural skills. It includes 180 tasks across six environments and finds that current agents often struggle to form robust reusable skills, with raw trajectory reuse often outperforming distilled skills.