@charliewarren: https://x.com/charliewarren/status/2062204573549490516
Summary
The article discusses the challenges of building AI native service companies, emphasizing that reducing variance in work quality is critical for trust and scale, more so than operating leverage alone.
View Cached Full Text
Cached at: 06/03/26, 09:55 PM
Trust in AI Native Service Companies
AI native service companies are startups rebuilding insurance carriers, law firms, accounting practices, and more with LLMs doing much of the work. The category is early. But the prize is the trillions of dollars in labor spend vs. the (now quaint) IT budgets of old.[1] Like many applications of AI, hype abounds.
The armchair caution for founders to date has focused on gross margins.[2] Can these companies grow revenue faster than COGS and compress cycle times? “Focus on scaling” is the refrain I hear most often.
Having now worked with these new kinds of startups, I would wager that operating leverage will be necessary yet insufficient to find product market fit.
AI native service companies will succeed based on whether they can reduce “variance.” By variance, I mean significant inconsistency in the quality of the work product delivered to the end customers. Poor quality work erodes trust and creates churn. And humans, not LLMs, are likely to drive the greatest variance for these startups.
For AI native service companies, variance is the enemy of trust. Model evals alone will not suffice.
How to Start
Jake Saper at Emergence and Julien Bek from Sequoia have written the best category overviews so far here and here, respectively. Both focus on the broader opportunity at hand and pitfalls running these businesses post Seed.
But how does a founder start one of these companies? I’ve had the pleasure of working with some great YC founders building in everything from healthcare services, to regulatory compliance, to devrel. The canonical startup advice does not always map perfectly to these companies. So I recorded a YC Startup School video specially for founders contemplating AI native services:
Y Combinator@ycombinator·5hSome of the biggest companies of the next decade won’t be software businesses. They’ll be services companies like insurance carriers, law firms, and tax practices rebuilt from scratch with AI doing most of the work.
In this episode of Startup School, YC Visiting PartnerShow more203730729K
I cover picking a market, forming the team, building the product, serving customers, the P&L, and whether you should buy a business. It’s a 101 course for builders exploring this space.
Spoiler alert: buying a company almost never works. Apologies to the PE cosplay crowd.
Variance vs. Trust
Starting an AI native service company comes with plenty of traps, many of which I detail in the video. Variance is the most important pitfall to avoid, yet the least understood.
To be sure, operating leverage is required. The last wave of “tech-enabled services” showed that technology plastered on a services business does not a priori create a software P&L – or valuation. Let’s not forget the painful experiments from the last decade: Compass (real estate agents), ScaleFactor (accounting), WeWork (coworking…and preschools), and Katerra (construction!), among others. These implosions had a variety of causes, but put simply: the businesses did not scale. Schadenfreude aside, the lesson is not simply that margins matter. It is that services businesses fail when they cannot reliably produce quality at scale.
The WeGrow preschool. Coworking, but for toddlers. It did not scale.
The WeGrow preschool. Coworking, but for toddlers. It did not scale.
AI may finally change the operating leverage equation, allowing companies to grow revenue faster than COGS in much bigger labor TAMs. But LLMs do not eliminate operational variance. If anything, the variance problem might be worse.
The counterintuitive reason for the persistence of variance: humans are still in the loop.
AI native services are not pure technology products. They are production systems that combine LLMs, internal workflows, customer inputs, exception handling, and human-in-the-loop decisions. Humans review outputs from the company’s internal product, resolve edge cases, manage customer-specific requests, and decide when a deliverable “looks done.” That judgment is the raison d’être of these businesses. And it is inconsistent precisely because it is human.
Some quality management and operations research is instructive here. In Out of the Crisis (1982), W. Edwards Deming distinguishes between “special cause” vs. “common cause” variation in manufacturing.[3] Common cause variation is the expected noise: random, small differences in product output that come from how the production process is designed. However, special cause variation is different in kind: it is the specific breakdown or defect that causes the customer to lose trust in the quality of the finished product. Founders need this distinction because not all errors deserve the same response.
For example, an AI-native law firm might produce flawless customer contracts most of the time, but miss an important indemnity term that stalls the largest deal at the end of the quarter. Meanwhile, an AI-native accounting firm might close the books correctly for a startup three quarters in a row, then misclassify the deferred revenue at year end before the first audit. These issues are not just “variance” in the abstract. The errors cause customers to lose trust in the AI native service and, eventually, churn.
If “quality management” doesn’t excite you like this guy, these are not the startups for you. Deming pictured in the 1950s while lecturing in Japan.
If “quality management” doesn’t excite you like this guy, these are not the startups for you. Deming pictured in the 1950s while lecturing in Japan.
Reducing special cause variation should be the obsession of AI native service companies.
Customers do not care that the average output quality is good, or that the company is using AI to deliver the results. They care about whether each output is correct, every time. Trust is a function of this output consistency. Conversely, lack of trust drives churn.[4]
Human Judgement & Process Evals
So how can an AI native service company reduce variance and maintain trust?
LLMs do create some unique challenges through their non deterministic outputs. All of the normal tactics and evals are important here. The models will continue to improve, too.
In general, the variance culprit will not be the models but rather the humans. The solution: AI native service companies need “process evals.”
Model evals like SWE-bench tell one how a model performs on bounded coding tasks. Harvey’s recent Legal Agent Benchmark extends the concept, testing whether legal agents can navigate complex client matters and produce reviewable work products. But even these benchmarks judge the technology outputs alone.
AI native service companies need to push this approach even further: process evals for the entire system, including the humans in the loop. These evals should measure the end-to-end delivery system, including customer intake, model output, handoffs, human review, exception handling, QA, and customer feedback on the final work product. They should track reviewer disagreement, exception rates, rework rates, escalation details, customer-reported errors, and whether certain humans or handoffs create repeatable failure modes. Once the process eval is in place, the company can refine their approach and reduce variance. The company also needs to build great products to attract and retain the best humans-in-the-loop. This is no small feat, either. These are early startup employees, not literal cogs in a wheel.
Creating these internal process evals will become the core IP of these businesses as the models get better yet us humans stay, well, the same. One could even imagine an entire startup that does third-party benchmarking of AI native services in particular industries.
Silicon Valley, meet the staid world of Six Sigma.
To build trust and reduce variance, founders need to make the end-to-end process quantifiable and constantly improving. The winners will measure the humans, not just the models.
Reposted from https://bearing.substack.com/
Notes
-
As I’ve written, there are plenty of untapped vertical AI markets in deskless industries where founders can build massive businesses. Not everything should be AI services
-
I do think founders should familiarize themselves with Little’s Law on throughput and cycle time as well as Kingman’s formula on utilization. On the latter, founders will be pressured to max out the humans in the loop to prove operating leverage, but counterintuitively slack in the system is the buffer against wait times growing exponentially. These will be hard learned lessons for many looking to impress VCs with top line growth.
-
Deming was heavily influenced by Walter Shewhart’s 1931 work, Economic Control of Quality of Manufactured Product. I could mention more here, but it risks turning this essay into Good Will Hunting’s bar scene with the ponytailed graduate student.
-
Measuring churn in these businesses is somewhat novel. Is there a gross dollar retention equivalent if the work is project based without a retainer? We shall see. Net dollar retention will be important, too.
Similar Articles
@rhythmrg: https://x.com/rhythmrg/status/2066561780495896785
The article argues that enterprises should post-train their own custom AI models for mission-critical, high-volume use cases to achieve differentiation, cost savings, and control over tradeoffs, rather than relying solely on general frontier models.
@lukepierceops: https://x.com/lukepierceops/status/2091858914757190013
The article argues that AI adoption in enterprises often fails because companies focus on buying tools rather than redesigning workflows, leading to low ROI and minimal impact. It emphasizes the need for architectural changes to achieve real success.
@sgurumur: https://x.com/sgurumur/status/2057916874546090132
An op-ed discussing the gap between AI code generation and production-grade systems, emphasizing that human judgment and domain expertise remain critical for orchestrating interconnected decision loops in complex domains.
@startupideaspod: https://x.com/startupideaspod/status/2087661842231664780
Allie Miller shares her playbook for building an AI-native company with a 34-agent workforce, using a three-word prompt and a management style focused on deciding rather than delegating.
@djfarrelly: https://x.com/djfarrelly/status/2052779234234380479
The article argues that AI agent development should rely on stable execution primitives rather than rigid frameworks, which frequently change with emerging orchestration patterns. It emphasizes durable steps, persistent state, parallel coordination, event-driven flow, and observability to prevent costly rewrites as best practices evolve.