@ArizePhoenix: You can use PXI to run an experiment directly from Phoenix! Here's one that tests the system prompt vs. schema-aware pr…
Summary
Arize Phoenix demonstrates using PXI to run an experiment comparing system prompt vs schema-aware prompt with a programmatic code evaluator, avoiding the need for an LLM judge.
View Cached Full Text
Cached at: 07/27/26, 03:56 PM
You can use PXI to run an experiment directly from Phoenix! Here’s one that tests the system prompt vs. schema-aware prompt, same model, graded by a code evaluator — no LLM judge needed when the check is programmatic.
TIL: “When an eval fails everything, suspect the eval first.” https://t.co/5X5Yw9oGFR
Similar Articles
@ArizePhoenix: PXI (Phoenix Intelligence) now runs in your terminal! You can now use PXI without leaving your terminal. It's the same …
PXI (Phoenix Intelligence), the AI agent previously only available in-browser, is now available as an interactive chat in your terminal via the CLI package @arizeai/phoenix-cli.
@ArizePhoenix: This week we make moving from questions to answers faster. • Smarter PXI workflows complete multi-step UI tasks now run…
Arize enhances PXI workflows by integrating a JavaScript sandbox to automate multi-step UI tasks, enabling the assistant to control the UI via code for faster results.
@ArizePhoenix: This week in Phoenix - feedback gets more visible and the agent gets more capable: Server-side bash for PXI subagents (…
This week's Phoenix update adds server-side bash for PXI subagents with sandboxed execution and built-in GraphQL access, improving feedback visibility and agent capabilities.
@ArizePhoenix: Phoenix has an agent built into it now! PXI can help you find the crucial traces you should actually be reading. Short …
Phoenix has a built-in Pixie assistant that helps users quickly filter out silent failure traces where agent spans have errors but model responses are normal, greatly improving trace reading efficiency.
@ArizePhoenix: Experiment Baselining and Charts When trying to determine if a new model is up to the task, you need to factor in many …
Phoenix now includes customizable experiment charts and baselining, allowing users to compare models along performance, latency, tokens, and cost dimensions, and set baselines for preferred models.