Tag
Course notes for Chapter 3 of AI Engineering, covering evaluation methodologies (Evals) in depth: starting from the challenges posed by open-ended outputs and black-box models, it introduces language modeling metrics such as cross-entropy, perplexity, BPC/BPB, and compares two mainstream evaluation approaches — AI-as-a-judge and comparative evaluation.