Tag
An unofficial Jev plugin for coding agents has been released, providing best practices, an API reference, and links to over 150 community projects, with evaluations demonstrating a 96% pass rate in coding tasks compared to other plugins.
Palantir's AI platform architecture for secure organizations highlights ontology-based tools, model agnosticism, and rigorous logging and evaluations for agent systems.
The tweet emphasizes the need for companies to own their AI intelligence stack, including custom models, harnesses, and evaluations, rather than renting, with a quote from Gabe Pereyra discussing challenges at Harvey.
OpenAI outlines priorities and principles for effective third party assessments to enhance AI safety, emphasizing independent scrutiny, shared standards, and deep access for rigorous evaluation.
The article reveals that misconfigurations at Irregular, an AI security company connected to Effective Altruism, led to AI agents accidentally accessing real systems during cybersecurity evaluations, raising concerns about AI safety testing practices.
The article announces the release of eight free, open-source AI agent skills with evaluations, highlighting top picks like an orchestrator using multiple AI models, a reader simulator, and a project analyzer.
Cognition has introduced SWE-2, a coding model that matches recent frontier models on leading evaluations while reducing costs by up to 70%, now accessible through Synara with Devin subscriptions.
A tweet highlighting that AI involves more than just ChatGPT, emphasizing the importance of integrating models, RAG, agents, memory, security, and evaluations.
Ant's Ling team has released Ling-3.0-flash-Fin, a finance-enhanced AI model designed for financial workflows, with free access on OpenRouter for one month and plans to open-source the weights.
A GitHub repository featuring over 100 open-source AI agent apps and skills with proper evaluations, compatible with multiple LLMs including Claude, Gemini, GPT, and open-source models.
This paper critiques existing evaluations of AI moral reasoning for focusing on moral values while overlooking moral norms, and proposes a research agenda to develop standardized methods and datasets for assessing normative reasoning in large language models.
This article outlines a 12-step roadmap for AI Agent Engineers in 2026, focusing on seven interconnected pillars like context, tools, and memory, with Claude-based workflows to build reliable production agents.
Future AGI is an open-source platform combining evaluations, tracing, simulations, guardrails, and optimization to help teams ship self-improving AI agents, with a nightly release available for early testing.
Miles Brundage highlights excellent work from the AI Security Institute investigating the impact of test-time compute budgets for frontier AI model evaluations, with praise from Noam Brown.
This tweet summarizes key takeaways from a detailed writeup on building complex cybersecurity evaluations for AI agents, covering long-horizon tasks, sourced benchmarks, varying difficulty levels, outcome verification challenges, and QA-based progress testing.
This article argues that AI agents lacking proper evaluations are not yet viable products, emphasizing the need for rigorous testing and benchmarking in AI development.
Summary of three key takeaways from conversations with leading AI researchers at CAISconf, covering the importance of evaluations for AI agents, the trade-offs between industry and academia, and a novel pedagogical RL approach.
This paper demonstrates that allowing attackers to strategically choose when to attack (attack selection) in agentic AI control evaluations significantly reduces measured safety, suggesting that current evaluations may overestimate safety against selective attackers.
Discusses the need for evolving AI evaluation benchmarks through difficulty, quality, and diversity refinement, citing examples like MMLU-Pro, MMLU-Redux, BIG-Bench Extra Hard, RealMath, MathArena, and DatBench.
A developer asks for recommendations for open-source alternatives to LangSmith for tracing, evaluations, and debugging agent workflows, citing restrictive paywalls.