evals

Tag

Cards List
#evals

Training a local LLM using CPT and RAG (with evals)

Reddit r/ArtificialInteligence ↗ · 13h ago

The article describes a project with experiments on training a local LLM using continued pretraining (CPT) and RAG for domain-specific knowledge, featuring comprehensive evaluations and findings.

0 favorites 0 likes
#evals

Interesting disruptive use cases of Jev for me. Add yours.

Reddit r/AI_Agents ↗ · 3d ago

The article presents several disruptive use cases for Jev, a tool that can be applied to AI evals, voice AI, workflows, synthetic personas, and more, providing fast and quantifiable solutions for various AI tasks.

0 favorites 0 likes
#evals

I stopped using 100% as the success criterion for my coding-agent evals

Reddit r/AI_Agents ↗ · 6d ago

The author revised the success criterion for coding-agent evaluations by moving from requiring three perfect 88/88 runs to a criterion where each case must pass at least 2 of 3 runs, reducing randomness impact and improving stability.

0 favorites 0 likes
#evals

@danshipper: WE NOW HAVE A HEAD OF EVALS the work he's doing is honestly gamechanging. cannot wait to show it to you

X AI KOLs Timeline ↗ · 2026-08-28 Cached

Mike Taylor has announced his new position as Head of Evals at @every, a role highlighted as game-changing for AI evaluations by @danshipper.

0 favorites 0 likes
#evals

AI Engineer Notebooks – free, framework-free RAG/agents/evals on Colab

Hacker News Top ↗ · 2026-08-27 Cached

A free, framework-free set of Colab notebooks for learning the applied LLM stack, including RAG, agents, and evals, targeting AI engineers.

1 favorites 1 likes
#evals

The loop is the product, not the model — two talks this month said it from opposite ends

Reddit r/AI_Agents ↗ · 2026-08-25

Two talks and a blog post argue that the feedback loop and harness engineering are more important than model weights for owning AI intelligence in production, highlighting context management and cost considerations.

0 favorites 0 likes
#evals

Hiring Agents Is the Easy Part (4 minute read)

TLDR AI ↗ · 2026-08-13 Cached

The article argues that the main constraint on AI agent adoption is not capability but verification, including how companies define quality, evaluate ongoing performance, and compound feedback. It explores challenges like tacit standards, company-specific evals, feedback ownership, and self-improving loops.

0 favorites 0 likes
#evals

@jerryjliu0: We're not Palantir, but we do think a lot about evals and hillclimbing w.r.t. document processing. If you have really h…

X AI KOLs Following ↗ · 2026-08-09 Cached

Jerry Liu promotes LlamaParse and LlamaAgents for large-scale document extraction, emphasizing LLM evals and hillclimbing for accuracy and cost. He also connects FDE work with evals and RL environments.

0 favorites 0 likes
#evals

@jerryjliu0: The future of FDE work seems closely related with all work around evals/posttraining/RL envs. FDEs are effectively resp…

X AI KOLs Following ↗ · 2026-08-09 Cached

Jerry Liu shares thoughts on how forward-deployed engineer (FDE) work will shift toward defining goals, environments, and evals while automated optimization processes handle tactical implementation.

0 favorites 0 likes
#evals

Show HN: Product analytics (and evals) for agent sessions on your MCP

Hacker News Top ↗ · 2026-08-03 Cached

Armature is a new product analytics tool for AI agent sessions, capturing user intent, agent thinking, and success scores via an SDK for MCP servers, with automatic PII redaction.

0 favorites 0 likes
#evals

@omarsar0: recommended reading. the opportunity on harness engineering alone is hard to even measure. if you are an AI builder, ha…

X AI KOLs Following ↗ · 2026-08-01 Cached

Omar recommends reading Chamath Palihapitiya's AI investing guide, emphasizing the importance of harness engineering and evals for AI builders.

0 favorites 0 likes
#evals

Agent Behavior (Website)

TLDR AI ↗ · 2026-07-31 Cached

Agent Behavior is a format for writing behavior specs for AI agents in Markdown, enabling teams to define, review, and evaluate expected agent conduct across interactions.

0 favorites 0 likes
#evals

@OpenAI: We hope these experiments serve as a reminder that evals rarely measure models in isolation—they also measure a bundle …

X AI KOLs ↗ · 2026-07-29

OpenAI reminds developers that eval results depend on API settings and harness design, recommending the Responses API, retaining reasoning, and using compaction for best performance.

0 favorites 0 likes
#evals

@jxmnop: nothing against pangram, but you guys do realise you could just hook codex up to the pangram api and have it keep rewri…

X AI KOLs Following ↗ · 2026-07-26 Cached

A user comments that you could hook Codex up to the Pangram API to rewrite text until it passes a human evaluation.

0 favorites 0 likes
#evals

Building agents taught me the model is rarely the problem. What's your hard-won lesson?

Reddit r/AI_Agents ↗ · 2026-07-20

A developer shares hard-won lessons from building AI agents: focusing on tool design over model choice, using small loops instead of giant prompts, logging agent context, adding guardrails early, and creating small evals to catch bugs.

0 favorites 0 likes
#evals

@_philschmid: At the @aiDotEngineer World's Fair, I gave a talk on why vibe-checking agent skills breaks in production and how to bui…

X AI KOLs Following ↗ · 2026-07-20 Cached

At the AI Engineer World's Fair, Phil Schmid gave a talk on why vibe-checking agent skills break in production and how to build reliable automated evals using negative test cases, skill limits, and ablation tests.

0 favorites 0 likes
#evals

Agent failures should become evals, not just traces

Reddit r/AI_Agents ↗ · 2026-07-20

Advocates for treating agent failures as evaluation benchmarks rather than just trace logs, emphasizing the need for systematic testing of AI agent behaviors.

0 favorites 0 likes
#evals

oqoqo

Product Hunt ↗ · 2026-07-17

oqoqo is a tool for building evals and custom benchmarks for real-world tasks.

0 favorites 0 likes
#evals

What developers actually pick for agent reliability: LangSmith, Langfuse, Phoenix, Braintrust and Galileo, mapped across four layers.

Reddit r/AI_Agents ↗ · 2026-07-10

The article compares popular developer tools for agent reliability across four layers: tracing/evals, runtime guardrails, and gateway. It finds that no single open-source tool covers all layers, and most developers use a combination.

0 favorites 0 likes
#evals

Why Alignment Evals Need Calibration (8 minute read)

TLDR AI ↗ · 2026-07-08 Cached

This article argues that alignment evaluations (evals) need to be properly calibrated to be meaningful, discussing common pitfalls and techniques for improving calibration.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback