@ArizePhoenix: Experiment Baselining and Charts When trying to determine if a new model is up to the task, you need to factor in many …
Summary
Phoenix now includes customizable experiment charts and baselining, allowing users to compare models along performance, latency, tokens, and cost dimensions, and set baselines for preferred models.
View Cached Full Text
Cached at: 07/11/26, 01:30 PM
Experiment Baselining and Charts
When trying to determine if a new model is up to the task, you need to factor in many dimensions.
performance - measured by evals latency - is the model fast enough to give you the right UX tokens - how chatty is the model to achieve the result cost - can my team afford to use this model
That’s why Phoenix now comes with customizable experiment charts and baselining. Compare experiments along the dimensions you care about and set a baseline once you’ve found a model you like.
Similar Articles
@ArizePhoenix: Phoenix 17.7.0 makes your token usage legible. New token detail charts break prompt + completion tokens into their part…
Phoenix 17.7.0 adds token detail charts that break down prompt and completion tokens into subcategories, with pan, zoom, and live-stream capabilities for better observability of AI model token usage.
@ArizePhoenix: • Faster trace analysis: use natural-language filters, one-click chart zoom, and dedicated annotation columns. • Expand…
Arize Phoenix announces updates including faster trace analysis with natural-language filters and an expanded REST API for managing retention assignments and model providers.
@ArizePhoenix: Our favorite tools are the ones that have maximum customizability. Last week we added customizable charts, command K, a…
Arize Phoenix announces new customization features for its AI agent monitoring platform, including customizable charts, command K, recent searches, and custom column ordering.
@ArizePhoenix: You can use PXI to run an experiment directly from Phoenix! Here's one that tests the system prompt vs. schema-aware pr…
Arize Phoenix demonstrates using PXI to run an experiment comparing system prompt vs schema-aware prompt with a programmatic code evaluator, avoiding the need for an LLM judge.
ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats
ChartArena is a comprehensive bilingual benchmark for chart parsing that evaluates models across eight chart families and three visual scenarios (digital, printed, hand-drawn), using a human-agent annotation pipeline and format-agnostic evaluation. Evaluations of 26 MLLMs reveal that while proprietary models lead overall, open-source models are catching up, and diagrammatic structures and hand-drawn scenarios remain challenging.