Tag
PHITSBench is an execution-scored benchmark for evaluating AI models on generating PHITS radiation-transport input files from natural language, covering editing, repair, and full generation tasks. Experiments with GPT-5.4 show that while domain knowledge improves performance, significant challenges remain in correctly configuring physical observables.