Evaluation Methodologies (Evals) in AI Engineering

Reddit r/ArtificialInteligence News

Summary

Course notes for Chapter 3 of AI Engineering, covering evaluation methodologies (Evals) in depth: starting from the challenges posed by open-ended outputs and black-box models, it introduces language modeling metrics such as cross-entropy, perplexity, BPC/BPB, and compares two mainstream evaluation approaches — AI-as-a-judge and comparative evaluation.

No content available
Original Article
View Cached Full Text

Cached at: 10/03/26, 06:57 AM

# AI Engineering Chapter 3: Evaluation Methodologies (Evals) **TL;DR:** The more capable and open-ended foundation models become, the harder they are to evaluate; this chapter starts from evaluation challenges, covers language modeling metrics such as cross entropy, perplexity, bits per character, and bits per byte, and introduces the two mainstream evaluation approaches of AI as a judge and comparative evaluation. --- ## Why Evaluation Matters So Much The course opens by making a clear point: foundation models (FMs, for short) are growing increasingly complex, and we need better ways to understand their capabilities, limitations, and room for improvement. Evaluation is the key step in taking AI applications from "toy-level demos" to production—when an AI application has to be delivered to consumers and customers, it is essential to make sure the model fits the requirements. Quality control for AI outputs serves two main purposes: - **Reducing risk**: Identify where the system might fail, and build evaluations around those areas. - **Finding opportunities**: Once risks are identified, evaluations also help uncover opportunities for improvement and innovation in language models or foundation models. Another important fact: foundation models contain language model (LM) components internally, so language model metrics transfer directly to foundation models—though not all of them apply universally. ### Open-ended Outputs Make Evaluation Inherently Difficult Foundation model outputs are open-ended. If you ask a model to summarize a passage of text, it can generate any summary—but how do you define whether a summary is "good"? This isn't like a classification problem or a true/false question with a clear answer. A summary has to account for coherence, fluency, relevance, and many other dimensions. Precisely because foundation models are open-ended, building evaluations (evals) around them is so difficult—and that is exactly why evaluation research matters so much. ### AI as a Judge: A Subjective Approach to Evaluation A popular practice today is "AI as a judge"—using AI to evaluate AI. The basic workflow is: first obtain an output from the foundation model, then use an AI to evaluate that output. The key point is that AI as a judge is inherently a **subjective** form of evaluation, for three reasons: 1. The judge itself is an AI model. 2. The evaluation result is largely determined by the prompt given and the judge model's architecture—different people using different prompts, or different models as judges, will get different results. 3. Adjusting sampling parameters (such as temperature) can also cause the judge to produce different outputs. Therefore, AI as a judge is not a standardized form of evaluation. --- ## 1. Challenges in Evaluating Foundation Models ### 1. Smarter models are harder to evaluate Smarter models are, of course, harder to evaluate. And because of the open-ended nature of foundation models, we often don't have a specific ground truth. Without golden-standard annotations, you can't directly measure right or wrong—you have to figure out how to build evaluations yourself, and that is precisely the difficulty. ### 2. Foundation models are black boxes Foundation models are often treated as black boxes, because most people don't understand how foundation models or language models work internally. If you don't understand how something works, it is hard to build methods to evaluate it. ### 3. Benchmarks go stale quickly Benchmarks used to test foundation models are frequently superseded by new, better models within six months to a year. Once a model achieves a high score on a benchmark, that benchmark can no longer test where else the model might improve—its usefulness gradually dies out. For example: - The **GLUE** benchmark was replaced by **SuperGLUE** in just a few months; - **MMLU** (Massive Multitask Language Understanding) was upgraded to a stronger benchmark within just a year, because the old benchmark could no longer test the capabilities of emerging models. So constantly producing newer benchmarks is itself a challenge. ### 4. General-purpose models make it hard to choose evaluation domains We don't know what an LLM might do, so we can't determine what kind of benchmark to build for it, or in which domains to test the model. Because of this, LLM evaluation, while it has gained some attention, has far from received the emphasis it deserves. Leading labs are currently investing resources and funding to incentivize more people to pursue LLM evaluation research. ### 5. Word-of-mouth evaluation isn't enough Today, people often evaluate models by word of mouth or by casually trying them out. But that is not the right way to evaluate a model—our AI models are heading into production, so we need a very rigorous, systematic way to evaluate them. --- ## 2. Language Modeling Metrics: Understanding Evaluation from the Perspective of Language Models We need to understand language modeling metrics because foundation models contain language model components internally; once you understand language model metrics, you can better understand foundation model evaluation. The language models underlying foundation models encode statistical information about the training data distribution. The core task of a language model is to predict the next token—the better its ability to predict the next token, the better it performs. Correspondingly, the stronger the predictive ability, the lower the values of metrics like cross entropy and perplexity. Common metrics in the field of language modeling include: - **Cross entropy** - **Perplexity** - **Bits per character (BPC)** - **Bits per byte (BPB)** Each metric has its own importance. ### Entropy: How Much Information a Token Carries Entropy measures how much information a token carries on average. The higher the entropy, the more information the token carries, and the more bits are needed to represent it. The course illustrates this with an example of grid positions: - **Splitting the grid in half** (top half / bottom half): there are only two possible states, representable with one bit—0 or 1. The position information here contains only one bit. - **Further subdividing the grid** (e.g., into four regions: 1, 2, 3, 4): now I know a more specific position—bottom right, directly below, bottom left, and so on. Binary encoding requires four representations: 00, 01, 10, 11—meaning two bits. You can see that in the second case, each token carries more information, but also requires more bits to represent it. Predicting such a very specific position is inherently harder; predicting "top or bottom" is much simpler. The conclusion is: **the more bits, the more information carried, and the harder it is for a language model to predict the next token**. Therefore, in this example, the subdivided case has higher entropy, and the halved case has lower entropy. ### Cross Entropy: How Hard Is It to Predict the Next Token Cross entropy measures **how difficult it is for a language model to predict what the next token is on a dataset**. The higher the cross entropy, the harder it is to predict the next token. A model's cross entropy on training data consists of two components: 1. **The predictability of the training data itself**: The training data also contains some randomness, so it has a certain amount of entropy. 2. **The difference between the model and the data distribution**: The model tries to match the distribution of the dataset, but it can never match it perfectly—only approximately. There is always some difference between the distribution captured by the language model and the true distribution of the dataset. In symbolic terms, the second component is the **KL divergence** between the true distribution of the training data and the distribution learned by the language model (denoted as Q). Therefore: > Cross entropy of the true dataset distribution relative to the language model's learned distribution = Entropy of the training data (self-entropy) + KL divergence between the two And KL divergence is always non-negative. Let H(P, Q) denote the model's cross entropy on the training data—this equation captures the two sources of cross entropy: the uncertainty inherent in the data itself, and the inevitable gap as the model approximates the true distribution. ### Bits Per Character (BPC) and Bits Per Byte (BPB) The unit of cross entropy can be expressed in bits. Suppose a language model has a cross entropy (CE, for short) of **6 bits**—this means the language model requires 6 bits to represent each token. The issue is: different models have different token granularities—for some models, a token might be a character; for others, a token might be a word, or two characters. - If a token contains two characters, then bits per character BPC = 6 ÷ 2 = **3**. But BPC has a clear flaw: **different character encoding schemes are not standardized**. - With **ASCII**, a character takes up roughly 7 bits; - With **UTF-8**, the number of bits per character can range from 8 to 32. For this reason, a more standardized metric is **bits per byte (BPB)**. BPB measures how many bits a language model needs to represent one byte of data. --- ## 3. Going Forward: Evaluation Methodologies and Comparative Evaluation The course also lays out two key directions for what follows: - **Specific evaluation methodologies**: Once the language modeling metrics are understood, moving on to evaluation methods that are closer to practice. - **Comparative evaluation**: Ranking models against each other by comparing them—this is a very popular way of evaluating AI today. --- ## Summary | Topic | Key Points | | --- | --- | | Why evaluate | Reduce risk, find opportunities, and take demo-level applications into production | | Open-ended outputs | No clear ground truth; must account for coherence, fluency, relevance, and other dimensions | | Main challenges | Smarter models are harder to evaluate; models are black boxes; benchmarks go stale quickly; evaluation domains are hard to pin down | | AI as a judge | Evaluating AI with AI, influenced by prompts, model architecture, and sampling parameters—essentially subjective | | Entropy | Measures the average information a token carries; more bits means harder to predict | | Cross entropy | Data entropy (self-entropy) + KL divergence between the model's distribution and the true distribution | | BPC / BPB | BPC varies with character encoding schemes and is not standardized; BPB is more standardized | Source: Evaluation Methodologies (Evals!) in AI Engineering (https://youtu.be/cT-H3oqbBIs)

Similar Articles

@yibie: https://x.com/yibie/status/2102619356874117594

X AI KOLs Timeline

This article discusses the importance of evaluations in AI systems, explains why traditional testing is insufficient, introduces three main types of evaluations, and provides implementation suggestions.

How evals drive the next chapter in AI for businesses

OpenAI Blog

OpenAI publishes a framework for business leaders on using AI evaluations (evals) to measure and improve AI system performance in organizational contexts, distinguishing between frontier evals for model development and contextual evals tailored to specific business workflows.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

Hugging Face Daily Papers

This paper introduces EvalCards, an operational framework that standardizes AI evaluation reporting by composing benchmark metadata, evaluation run data, and model metadata into a unified record with interpretive signals for reproducibility, completeness, provenance, risk, and score comparability. The authors deploy a monitoring tool across thousands of models and benchmarks, revealing systematic gaps in current reporting practices.

BEAMS: Benchmarking and Evaluating AI for Modeling and Simulation

arXiv cs.AI

The BEAMS Initiative presents a benchmark suite for evaluating AI tools in modeling and simulation, focusing on human-centered and responsible AI practices. Tests reveal variability across LLM-based engines, with better performance in qualitative tasks than causal reasoning.