Tag
Phoenix now supports Meta's Muse Spark 1.3, a model released quietly alongside others, and it deserves more attention than it received.
ArizePhoenix has released an update that automatically searches for duplicates, drafts GitHub issues with linked traces and spans, and uses a secure browser-based GitHub token for filing.
The article details the implementation of cost and token observability for LangChain applications using Arize Phoenix and OpenTelemetry, focusing on distinguishing between LLM tool calls, tool execution, and final responses to avoid collapsing metrics.
ArizePhoenix shows significant performance improvements on benchmarks like Terminal-Bench 4.0 and Humanity's Last Exam, with gains from previous versions such as Fable 5.
Arize Phoenix shipped four new instrumentors for AG2, Together AI, Cohere, and Ollama, expanding its AI observability and tracing integrations.
Arize Phoenix announces customizable visualizations for agent traces, enabling real-time tracking of cache hits, online eval degradations, and tool call errors in production.
Arize Phoenix announces new customization features for its AI agent monitoring platform, including customizable charts, command K, recent searches, and custom column ordering.
Phoenix now includes customizable experiment charts and baselining, allowing users to compare models along performance, latency, tokens, and cost dimensions, and set baselines for preferred models.
PXI (Phoenix Intelligence), the AI agent previously only available in-browser, is now available as an interactive chat in your terminal via the CLI package @arizeai/phoenix-cli.
This week's Phoenix update adds server-side bash for PXI subagents with sandboxed execution and built-in GraphQL access, improving feedback visibility and agent capabilities.
Arize Phoenix announces an experimental feature for PXI to spin off subagents, keeping the main context window lean during long investigations.
Arize Phoenix enables local-first, air-gapped observability for coding agents, allowing each agent to have its own traces, evals, and feedback loop for self-verification.
Arize Phoenix announces a free 2-hour evaluations workshop from the AI Engineer: Europe conference, led by head of DevRel Laurie Voss, covering manual data examination and built-in/custom evals.
The official TanStack AI OpenTelemetry support is now available, offering an open-source backend for traces, datasets, and replay to improve debuggability.
This article discusses best practices for LLM application development using Arize Phoenix, specifically highlighting the importance of using train/validation/test splits for honest evaluation and tracking regressions.