Tag
The paper introduces a six-dimension prompt-side structural complexity index to measure code-generation reliability in LLMs independently of functional correctness, evaluating 21 models on 5,000 Python prompts to identify nonmonotonic reliability regimes.
The article analyzes token cost breakdown in Claude Code deployments, revealing that only 14% of input tokens are user prompts, with the rest being configuration and context, and offers insights to reduce bills by 20-40%.
This paper analyzes real-world user queries about digital security and privacy asked to LLMs, categorizing them into nine topics and evaluating response quality and consistency across commercial and open-weight models.