Tag
The tweet highlights the essential role of evals in automating AI for enterprises, emphasizing that measurement is key to understanding and improving non-deterministic processes like agent performance.
Microsoft has unveiled a redesigned Copilot 'super app' that bundles chat, coding, and an AI assistant named Autopilot into a single interface, positioning it to be as influential as Office for enterprise productivity.
This preprint paper proposes a customizable router for AI coding agents in enterprises to optimize costs by intelligently routing requests, saving 14-21% of model spend annually for large companies.
A practical guide on using confidence scores in document extraction to balance automation and accuracy by setting thresholds for precision and recall.
Dropbox CTO Ali Dasdan shares lessons learned from deploying AI at company scale, focusing on rethinking workflows and measuring real impact.
The tweet discusses the importance of open vs. closed AI models and data sovereignty for enterprises and governments, particularly outside major AI superpowers, with links to a full conversation and a book.
The article features an interview with ElevenLabs' CEO discussing the company's AI voice technology, its $22 billion valuation, enterprise adoption, and future plans including IPO timing and ethical guidelines for AI interactions.
An announcement for talks at the Interrupt main stage covering AI agent building with LangChain, JPMorgan's agent assembly line, and scaling agentic AI at Omnicom.
James Luan, CTO of Zilliz, argues that data infrastructure becomes more critical as AI agents act on enterprise data, emphasizing the need for data quality, freshness, and proper permissions.
Stripe has launched the Knowledge AI Platform, a versatile AI agent system built to handle diverse non-coding knowledge work, from quick queries to complex projects, by connecting employees to over 1,000 internal tools.
Gabe Stengel discusses persistent agent memory as a major unsolved problem in enterprises, focusing on memory compaction to maintain context and coherence in AI interactions.
The author describes a three-week process of implementing AI in an AI town, shifting to a bottom-up approach where employees build their own agents using tools like Workbuddy and thincoder for internal deployment.
The article argues that in enterprise AI, the AI model is often the easiest part, with greater challenges in data trust, governance, and workflow integration, especially in regulated industries like healthcare.
EvidenT is a lightweight pipeline that enhances evidence groundedness and traceability in enterprise retrieval-augmented generation (RAG) systems, improving gold-source hit rate by 29% without model retraining.
The YC Early Access Network event on October 22 will bring together tech leaders from high-growth startups to Fortune 500 companies to meet leading enterprise AI startups, with spots filling up quickly.
Confido, an AI startup automating operations for consumer brands, raised $55 million in Series B funding led by Insight, with co-founders sharing their journey on the Founder Firesides podcast.
TechCrunch Disrupt 2026 will host five AI safety sessions aimed at founders, covering topics like agent security and enterprise deployment, with insights from Anthropic.
The article details a pilot project where a team tested client data for a RAG system, initially achieving only 40% accuracy due to issues like conflicting documents and poor retrieval. Through routing, data cleaning, and acronym mapping, accuracy improved to about 90%, highlighting the critical role of data readiness in successful AI transformation.
The post discusses challenges with handoffs in enterprise agentic AI systems when agents encounter exceptions, and it seeks advice on designing effective escalation paths to prevent issues from being ignored.
Codos launches a virtual Chief AI Officer that addresses context loss in Forward Deployed Engineers by deploying company-wide memory and automation agents, with preliminary benchmark scores showing high performance on EnterpriseRAG-Bench.