Tag
The article explores how AI labs and startups are using additional AI systems to monitor and control rogue AI agents, addressing the challenge of overseeing large-scale AI actions that exceed human review capabilities, while noting concerns about AI deception.
Phoenix now supports Meta's Muse Spark 1.3, a model released quietly alongside others, and it deserves more attention than it received.
Arize Phoenix now supports GLM-5.3, which builds on GLM-5.2 with improvements from post-training and training on scaled long-horizon environments using an open-source reinforcement learning framework.
Arize Phoenix announces updates including faster trace analysis with natural-language filters and an expanded REST API for managing retention assignments and model providers.
The article highlights that most teams use separate tools for AI tracking and introduces Respan as a unified solution to route, observe, and evaluate LLM calls through a single gateway.
OpenObserve launches an AI-native, open-source observability platform designed to trace AI agents and LLMs, offering detailed insights into performance, cost, and quality for developers.
A tweet discusses the growth of AI Dot Engineer in AI Observability and notes that Dynarize, a $14B observability company, has acquired an AI-native team.
The Linux Foundation has formally launched the Tokenomics Foundation, a vendor-neutral standards body focused on quantifying the true cost and value of AI, moving beyond the era of unchecked token spending.
Progress AI Observability is a product for tracing, evaluating, and improving AI agents in production.
Day 3 of Grafana Labs AI Week introduces new AI features for faster incident investigation and automated maintenance, part of a five-day event revealing agentic operations capabilities.
Oodle launches Agent Observability on Hacker News, offering agent traces at $10 per million spans to help AI-native teams detect silent failures and improve reliability.
Agnost AI is a platform that analyzes agent conversations to surface user feedback, detect failures, and generate automated fixes, helping teams improve their AI agents.
An article discussing the growing public fascination with watching AI agents perform tasks, and the question of what to call this phenomenon.
PostHog has open-sourced its all-in-one product analytics platform, offering tools like product analytics, web analytics, session replays, feature flags, and more. The platform is free to use with a generous free tier.
SaZabi is building an AI observability system that uses logs as the source of truth to automate debugging and issue resolution, aiming to bridge the gap between automated code and manual debugging.
Opik is an open-source platform for AI agent observability that goes beyond tracing to automatically diagnose failures, propose fixes, and verify them, closing the debugging loop without manual intervention.
Braintrust's Topics feature uses LLM summarization to make production agent traces tractable for clustering and classification at scale, inspired by Anthropic's Clio approach.
A developer built TracePilot, a lightweight zero-dependency npm SDK for AI observability to simplify debugging prompts in production, offering real-time latency, token costs, and error tracking.
Respan introduces an AI observability platform that automatically catches issues in traces, aiming to replace manual debugging for agent-based workflows.
Phoenix introduces Code Evaluators, allowing users to define evaluation strategies in Python or TypeScript directly in the UI, with server-side execution and composable scoring methods.