Tag
Introduces SciForma, a framework for generating scientific methodology diagrams with high structural fidelity, using multi-dimensional conjunctive preference optimization (M-DPO) and a structural inventory to ensure correctness across component, arrow, and text axes. The 9B model surpasses open-source baselines and GPT-Image-1.5 on benchmark evaluations.
LongWebBench is a benchmark for evaluating long-horizon webpage generation from both structural and functional perspectives, using VLM-based metrics and DOM-augmented agent-based pipelines. Experiments show current VLMs struggle with long-range coherence and executable interactions.
This article discusses the importance of faithfulness in LLM optimization, introducing a Structural Fidelity Score that measures drift across word overlap, constraint survival, and task-type match to ensure prompt optimization does not sacrifice intent.