semantic-reliability

Tag

Cards List
#semantic-reliability

When JSON Is Not Enough: Semantic Reliability of Schema-Constrained LLM Ordering Agents

arXiv cs.AI · yesterday Cached

The paper introduces OrderBench, a benchmark for restaurant ordering LLM agents that evaluates semantic reliability beyond schema validity, demonstrating that structured output modes can achieve perfect schema validity while still having high semantic error rates.

0 favorites 0 likes
← Back to home

Submit Feedback