@shao__meng: https://x.com/shao__meng/status/2105545533121294508
Summary
Stanford's Monica Lam team has released CS224V "Agentic AI," a research course focused on anti-hallucination and accountability. Using retrieval, declarative specifications, theorem provers, and rigorous evaluation, it turns unreliable LLMs into trustworthy agents, with projects accounting for 65% of the grade.
View Cached Full Text
Cached at: 10/02/26, 04:34 AM
Stanford CS224V “Agentic AI” Course: How Monica Lam’s Team Turns Hallucination-Prone LLMs into Trustworthy, Accountable Agents
Stanford CS224V “Agentic AI”
https://web.stanford.edu/class/cs224v/index.html
https://web.stanford.edu/class/cs224v/index.html
A project-centered agentic AI research course taught by Stanford professor Monica Lam (from the OVAL Lab, her NLP group). But its stance is the polar opposite of most “Agent crash courses” on the market: the entire course revolves around a single central question — how to turn LLMs, which are inherently unreliable and prone to hallucination, into trustworthy, accountable agents. The course believes such systems hold “extraordinary promise for accelerating scientific discovery and democratizing healthcare, legal, and educational services” — but only if reliability is solved first.
The first two lecture slide decks are already publicly available:
https://web.stanford.edu/class/cs224v/lectures_2026/l-introduction.pdf
https://web.stanford.edu/class/cs224v/lectures_2026/I-DeepResearch.pdf
The Course’s Technical Throughline: An “Anti-Hallucination” Knowledge Stack
Viewed together, the 15 lectures and the reading list form a complete technology stack built bottom-up, with each layer answering the same question: where does reliability come from?
1. Foundations: the LLM itself (Reading List §1) Only two papers are required: the original Transformer paper (Attention is All You Need) and Chain-of-Thought. The supporting readings are GPT-3, InstructGPT, LLaMA, DeepSeek-R1, plus tool use (MCP), distillation and test-time scaling (s1). This layer is just the foundation, not the focus.
2. Knowledge Curation: the STORM family (Lecture on 9/28: STORM/Co-STORM) Have the LLM iteratively search and read from multiple perspectives before writing — generating open-domain knowledge the way one would write a Wikipedia article. This is the Lam group’s signature work, and is often regarded as one of the academic origins of today’s “Deep Research” products.
3. Reliable Question Answering over Free Text (RAG) Retrieval (ColBERT, RankGPT) + generation + verification. The representative system, WikiChat (winner of Wikimedia’s 2024 Best Research Award of the year), emphasizes “verify first, then generate,” rather than generate-then-explain. This comes with an entire section of readings on factuality evaluation methodology — the course’s emphasis on evaluation runs throughout.
4. Agents over Structured and Heterogeneous Data This is the most distinctive part of the course. SUQL is “the first query language that combines information retrieval (natural language) with database queries (SQL).” SPINACH uses reactive agent workflows to tackle SPARQL knowledge-graph queries (15 billion facts in Wikidata) and has been deployed in production at Wikimedia. The core idea: whatever can be handed to a formal query engine should not be left to the LLM’s free play.
5. Formal Methods and Accountable AI (the course’s most unique hallmark) A lecture on 10/5 encodes constraints as SMT (satisfiability) formulas, letting a theorem prover serve as an “accountable decision-maker” that guarantees the LLM’s rationale is sound and complete. SatIR and VERDICT target high-stakes scenarios like clinical-trial matching. Related readings include SATLM, Logic-LM, Lang2LTL — i.e., the neuro-symbolic route in which “the LLM does the translation, and a symbolic solver does the reasoning.”
6. Task-Oriented Dialogue Agents Genie Worksheet is described as “the first declarative specification language for advanced task-oriented agents”: first use a formal dialogue-state representation to constrain the agent, then let the LLM handle semantic parsing and response generation — thereby scaling hallucination-free agents reproducibly across vertical domains.
7. Meta-Agents and Meta-Optimization (a look to the future) A lecture on 11/11 covers the anatomy of Coding Agents (Claude Code, Codex) — LLM + tools + harness — and meta-optimizing the harness itself. This section has the densest reading load (11 papers): ADAS, AFlow, DSPy, OPRO, TextGrad, AlphaEvolve, STOP, Darwin Gödel Machine, Dream-RSI — moving all the way from “prompt optimization” to “recursively self-improving agents,” tracking the 2025–2026 frontier closely.
There are also two lectures that are somewhat “surprising” yet very much in character for the Lam group: VLM transcription of historical manuscripts (CHURRO, done in collaboration with historians) and training LLMs (instruction-following models and training data).
The Project: The Real Star of This Course (65% of the Grade)
• Teams of two, running the whole quarter, meeting weekly with mentors (from Stanford or external partner institutions). The writing output is benchmarked against a small research paper (ACL format suggested).
• A clear hard bar: the project cannot be something you produce by “prompting the problem statement into a vanilla LLM” — it must employ the course’s techniques, such as formal methods, retrieval, and knowledge reasoning.
• An open-source platform is provided: AOS (Agentic OS) — model-agnostic, with a semantic file system. Its components (CHURRO, SLIDERS, GRILL, DataSTORM, Genie Worksheets, VERDICT) serve both as students’ toolkit and as research objects in their own right, ripe for improvement.
• Topics come along two axes: domain-driven (e.g., healthcare, with OpenEvidence’s 44% accuracy on MedXpertQA as the benchmark target, spanning knowledge curation, data-driven discovery, clinical guidelines, and EHR question answering) or technology-driven (improving AOS components). The 2026 fall mentor proposals are heavily concentrated on medical genetics: scaling OMIM-style database curation, literature-grounded writing for rare diseases, genomic sequencing diagnostics (breaking through the roughly 30% diagnostic yield), RNA-Seq gene prioritization for triple-negative breast cancer, and more.
• The track record is convincing: papers produced by this course have inspired the commercial Deep Research products from OpenAI, Gemini, and Databricks; the WWKnowledge.org pilot served over 800,000 users. WikiChat, SUQL, SPINACH, Genie Worksheets, DataSTORM, VERDICT, and others have been published at mainstream venues such as EMNLP, ACL, NAACL, and COLM.
Organization and Assessment
• Classes meet Monday/Wednesday; attendance is mandatory. Prerequisites: NLP fundamentals (CS 224N or equivalent). Enrollment is capped to preserve the quality of project mentoring.
• Grading: participation 15% + assignments 20% + final project 65%. That weighting makes it obvious this is a research apprenticeship — the lectures are just scaffolding.
• The schedule is built around the project’s rhythm: around the midterm, students give pitches and defend proposals; at the end of the term, there are 60 lightning talks + a poster session + a paper.
What Makes This Course Special?
The most notable thing about CS224V is that, in the context of the 2026 “Agent gold rush,” it chose a counter-current path. Industry at large bets on emergent reliability through scale, alignment training, and end-to-end reinforcement learning. This course’s philosophy is systems engineering + formal constraints: treat the LLM as a powerful but unruly language/reasoning component, ground it with retrieval, constrain it with query languages and declarative specifications, guarantee it with theorem provers, and judge it with rigorous evaluation. The systems cited on the course page — SPINACH deployed at Wikimedia, WWKnowledge serving 800,000 users, VERDICT in clinical deployment — show that this route is not just paper aesthetics; it holds up in real-world settings.
The structure of the reading list also reveals the course’s pedagogy: the required readings (~12 papers) are almost entirely the Lam group’s own system papers, while the supporting readings span the whole field (from ColBERT to AlphaEvolve) — and “LLM evaluation” is broken out as an entire section of its own. What this course teaches is not just “how to build agents,” but “how to know that the agent you built is worth trusting.”
Who it’s for: people with solid NLP foundations who want to spend a full quarter doing agent research inside a real research ecosystem (not a toy environment). Those interested in the “neuro-symbolic + accountable AI + healthcare deployment” route will feel especially at home. Who it’s not for: engineering speedrunners who want to “build an Agent in three days” — 65% of the course is a research project, and pure prompt engineering is explicitly ruled out.
Similar Articles
@WangNextDoor2: Stanford CS146S: A Must-Take AI Programming Introductory Course https://heyuan110.com/zh/posts/ai/2026-02-24-stanford-cs146s-overview/…
Stanford University's new CS146S course systematically teaches AI programming (Vibe Coding), covering LLM principles, Agent architecture, MCP, etc. All resources are free and publicly available, marking AI programming as a formal engineering discipline.
@tan_maty: Oh my god, the AI Stanford course shared by the awesome @alisawuffles who starts at OpenAI next week — I found it! Must-see for beginners! I've already learned it (and lost my mind), come join me! I feel my English improving too! Stanford CS336: Language Mod…
Stanford CS336 aims to teach students how to build language models from scratch, with deep understanding of the full-stack design of data, systems, and models. The course videos are publicly available and suitable for AI beginners.
@0xcryptowizard: Stanford's latest course, worth checking out. About self-evolving AI, covering scaling laws, chain-of-thought, RLHF, reasoning models, etc. YouTube link: https://youtu.be/6YnLB0XbTnI?si=Dg-aSXbDxymA4U…
Recommending Stanford University's CS329A course, about self-improving AI agents, covering scaling laws, chain-of-thought, RLHF, reasoning models, and more.
@shao__meng: https://x.com/shao__meng/status/2102648813425135777
本文总结了来自 Hamel Husain 和 Shreya Shankar 的 AI Evals 课程的 8 个核心技能,旨在指导工程师和产品经理构建有效的 AI 评测系统,涵盖错误分析、评测器设计、校准和监控等步骤。
@yiliuai: One of my biggest takeaways at MIT is that the world is just a giant makeshift operation. I experienced it deeply again today. This semester I'm taking a course on AI Agents, which requires group projects (the professor said he prefers infra over app layer). Inspired by my previous project connecting an LLM to a vibrator, and because I'm interested in hardware, I originally wanted to build a protocol or middleware layer that allows AI agents to connect and control most hardware. Coincidentally, another MIT undergrad wanted to do something similar; he invited me to join after learning about my vibrator project. But after careful consideration, I abandoned this project and instead chose another more superficial app-layer project. (I don't care about grades anymore, and I couldn't find a project I was both interested in and saw promise for.) My reason for giving up on this direction I was initially very interested in: I believed it had neither commercial value nor any technical moat.
The author shares an experience in an MIT AI Agent course, reflecting on the reasons for abandoning a project aimed at connecting AI agents to hardware, and criticizes a competing project that won high praise from judges and VCs merely by integrating and packaging a simple robotic arm interface.