@undefinedKi: DoorDash just published the full structure behind Ask DoorDash, their new AI assistant. And it's the clearest picture o…

X AI KOLs Following News

Summary

DoorDash shared how they automated evaluation of their Ask DoorDash AI assistant, enabling 2,000 daily graded sessions and cutting test time from six hours to 20 minutes while halving error rates.

DoorDash just published the full structure behind Ask DoorDash, their new AI assistant. And it's the clearest picture of the AI job nobody advertises Ask DoorDash is the assistant inside their app. You type a question, it answers, finds things and places the order. Millions of conversations so far. Their problem at the start: no way to tell whether a change made it better. Checking quality meant employees writing feedback by hand, around one note a day. A full test round took over six hours, so it almost never ran. So they built something that grades the assistant for them: 1. Write down what a good answer looks like, as explicit checks. 2. Rebuild real sessions from logs, so the grader sees what actually happened. 3. Replay them with simulated users and frozen tool responses, so a rerun measures your change and nothing else. 4. Let a model do the grading, after calibrating it against human labels. Now it grades 2,000 sessions a day, a full test round takes 20 minutes, and error rates dropped nearly by half before the national launch. Nobody at DoorDash has eval engineer on a business card. Someone still writes the rubric and calibrates the judge. That job is arriving at every company that ships a model, food delivery included. Bookmark this
Original Article
View Cached Full Text

Cached at: 08/06/26, 12:34 AM

DoorDash just published the full structure behind Ask DoorDash, their new AI assistant. And it’s the clearest picture of the AI job nobody advertises

Ask DoorDash is the assistant inside their app. You type a question, it answers, finds things and places the order. Millions of conversations so far.

Their problem at the start: no way to tell whether a change made it better. Checking quality meant employees writing feedback by hand, around one note a day. A full test round took over six hours, so it almost never ran.

So they built something that grades the assistant for them:

  1. Write down what a good answer looks like, as explicit checks.

  2. Rebuild real sessions from logs, so the grader sees what actually happened.

  3. Replay them with simulated users and frozen tool responses, so a rerun measures your change and nothing else.

  4. Let a model do the grading, after calibrating it against human labels.

Now it grades 2,000 sessions a day, a full test round takes 20 minutes, and error rates dropped nearly by half before the national launch.

Nobody at DoorDash has eval engineer on a business card. Someone still writes the rubric and calibrates the judge. That job is arriving at every company that ships a model, food delivery included.

Bookmark this

Similar Articles

Q&A with DoorDash’s CPO, Mariana Garavaglia

OpenAI Blog

DoorDash's Chief People Officer Mariana Garavaglia discusses how the company is scaling AI adoption across its organization to empower employees, measuring AI literacy through adoption metrics and frequency of use, and using AI tools to democratize automation beyond technical teams.

Empowering teams to unlock insights faster at OpenAI

OpenAI Blog

OpenAI has developed an internal research assistant that combines dashboards with a conversational GPT-5 interface to help teams analyze millions of support tickets and generate insights in minutes instead of weeks. The tool democratizes data analysis across teams, allowing non-technical users to ask questions in plain language and get actionable reports on product feedback, customer sentiment, and trends.