@MaxForAI: You'd be hard-pressed to find a better eval resource library. If you're interested in eval, these are what you should read. Thanks to @xdotli for sharing.
Summary
Share a curated AI evaluation (evals) resource library, including high-quality blogs, podcasts, papers, and projects, compiled by Xiangyi Li.
View Cached Full Text
Cached at: 06/24/26, 08:29 PM
It’s hard to find a better eval resource library than this one.
If you’re interested in eval, these are what you should read.
Thanks to @xdotli for sharing
Xiangyi Li (@xdotli): sharing my personal library on evals 1/n
i put together the highest quality blogs, podcasts, papers, and projects on evals. additions are welcome!
The Unsloth team is terrifying.
They took China’s top open-source model, GLM 5.2, and optimized it using an extreme technique called 1-bit compression, then converted it into a lightweight GGUF format.
This means you can run GLM 5.2 entirely locally (on a 256GB Mac Studio) with no internet connection or external server needed, and at an impressive speed of 21 tokens per second (human speech is about 10-20 tokens per second).
And they didn’t stop there — they also livestreamed a showdown, pitting this compressed local model against the world’s most powerful and expensive paid cloud models: Claude 4.8 Opus and GPT-5.5.
What shocked developers in the comments the most was that this local model actually traded blows with multi-billion-dollar server clusters, delivering intelligent and precise answers that rivaled these closed-source giants.
The real winner today isn’t a specific model — it’s the concept of local inference:
From now on, your data stays 100% safe on your device
Your API bill is zero dollars, with intelligence on par with the best US companies
What a crazy but revolutionary approach.
Similar Articles
@xdotli: sharing my personal library on evals 1/n i put together the highest quality blogs, podcasts, papers, and projects on ev…
A Twitter thread sharing a curated personal library of high-quality blogs, podcasts, papers, and projects on AI evaluations (evals), inviting additions.
@yibie: https://x.com/yibie/status/2102619356874117594
This article discusses the importance of evaluations in AI systems, explains why traditional testing is insufficient, introduces three main types of evaluations, and provides implementation suggestions.
@pauliusztin_: Every day, 100+ people ask me, "How can I learn AI evals?" I copy-paste these 11 links (every time): 1. AI evals & obse…
A curated list of 11 links shared daily to help people learn AI evaluation techniques, covering evals, observability, LLM-as-judge, and agent evaluation.
@Jackywxsz: https://x.com/Jackywxsz/status/2070743217721770316
The article recommends 5 high-quality AI information sources (AI Hot, BestBlogs.dev, AIGC Weekly, TW93 Blog, The Road to AGI), and shares the author's information acquisition logic based on attention management, emphasizing depth rather than chasing trends.
@zhizhuxia22: I've compiled yet another collection of top AI bloggers 1) @lxfater (Iron Man) Practical AI startup experience, multiple open-source projects 2) @xiaohu (Xiaohu) Cutting-edge AI trends, explained in an easy-to-understand way 3) @vista8 (Xiangyang Qiaomu) Plenty of practical tutorials, frequent tool sharing 4) @oran_ge (Orange AI)…
This is a post sharing a collection of top AI bloggers and an Agent Skills website, aimed at helping readers find AI learning and tool resources that suit them.