production-monitoring

Tag

Cards List
#production-monitoring

I ran the same prompt against our agent every week for a quarter and watched the answers drift until they broke our policy [D]

Reddit r/MachineLearning ↗ · yesterday

The author describes running weekly audits on a production AI agent, observing that its responses gradually drifted and eventually violated policy without model updates, emphasizing the need for continuous monitoring in real-world AI deployments.

0 favorites 0 likes
#production-monitoring

@LangChain: On September 30th, learn about the key elements of online evals and see how teams can use Tuned Evaluators in LangSmith…

X AI KOLs Following ↗ · 4d ago Cached

LangChain announces an event on September 30th focused on online evaluations, demonstrating how to use Tuned Evaluators in LangSmith to automatically analyze AI agent interactions and provide feedback in production.

0 favorites 0 likes
#production-monitoring

@KlausCodes: At @Simulithic we simulate user-behavior While building it we realized we were monitoring production every single time …

X AI KOLs Timeline ↗ · 2026-09-17 Cached

Simulithic simulates user behavior to monitor production and is making this tool available for early users, with an invitation to schedule a call.

0 favorites 0 likes
#production-monitoring

Built a verification layer that checks Postgres after your agent claims success, instead of trusting the 200 OK

Reddit r/AI_Agents ↗ · 2026-09-17

The author built synathic, an SDK that verifies Postgres database writes after agent functions report success to ensure data integrity, addressing common issues where agents claim success without actual changes.

0 favorites 0 likes
#production-monitoring

what do u actually check when every span is green but the agent still did the wrong thing?

Reddit r/AI_Agents ↗ · 2026-09-10

A user discusses strategies to debug AI agent systems in production where all indicators show success but outcomes are incorrect, seeking community advice on evidence and methods for diagnosis.

0 favorites 0 likes
#production-monitoring

I analyzed 31,352 hourly LLM benchmark scores: within-day variation was 2.8 points, while between-day variation was 8.4 [P]

Reddit r/MachineLearning ↗ · 2026-08-29

An analysis of 31,352 hourly LLM benchmark scores shows between-day variation is about three times greater than within-day variation, emphasizing the importance of continuous monitoring for performance drift, leading to the creation of the AIStupidLevel system.

0 favorites 0 likes
#production-monitoring

@cursor_ai: We’re excited to welcome the Firetiger team to Cursor! Together, we're building agents that can follow their work into …

X AI KOLs Timeline ↗ · 2026-08-13 Cached

Firetiger, a startup building agents that monitor and fix production software, is joining Cursor. The team will help Cursor build long-running autonomous agents that can ship code, observe its behavior in production, and respond to issues.

0 favorites 0 likes
#production-monitoring

How are you evaluating AI features in production?

Reddit r/AI_Agents ↗ · 2026-06-23

A discussion on the methodologies and challenges involved in evaluating AI features once they are deployed in production environments.

0 favorites 0 likes
#production-monitoring

@hwchase17: Detecting issues in production agent traces is hard. You have to do it cheaply (because of volume) but also accurately …

X AI KOLs Following ↗ · 2026-06-15

Harrison Chase announces a post-trained model for detecting issues in production agent traces, claiming SOTA accuracy at 10-100x cheaper rates than frontier models.

0 favorites 0 likes
#production-monitoring

@_avichawla: Claude Code without this new tool is like Git without GitHub. Claude Code stops at the boundary of your terminal. - It …

X AI KOLs Timeline ↗ · 2026-06-08 Cached

CodeRabbit Agent integrates with Claude Code and Slack to bridge operational and institutional memory gaps, enabling automatic incident tracing, root cause analysis, and documentation without switching between dashboards.

0 favorites 0 likes
#production-monitoring

How are teams handling prompt QA at scale?

Reddit r/AI_Agents ↗ · 2026-05-20

A practitioner at a company handling ~40k conversations/month describes the bottleneck of manual prompt QA and asks how teams are using automated systems to detect regressions and user frustration in production.

0 favorites 0 likes
#production-monitoring

@dabit3: Most coding agents still live in the “write code” part of the SDLC. The next era of AI software development is moving a…

X AI KOLs Following ↗ · 2026-05-18 Cached

The next era of AI software development moves coding agents into production; Cognition introduces Devin Auto-Triage for automated incident response and PR generation.

0 favorites 0 likes
#production-monitoring

@ds3638: Evals are dead. Or more precisely: traditional eval-driven development doesn’t scale. Static evals were useful when age…

X AI KOLs Timeline ↗ · 2026-05-14

Traditional eval-driven development doesn't scale for long-running autonomous agents; observability-driven development—with tight guardrails, production trajectory collection, and behavior clustering—is becoming the foundation for prod-ready AI systems.

1 favorites 1 likes
← Back to home

Submit Feedback