scientific-benchmark

Tag

Cards List
#scientific-benchmark

SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models

arXiv cs.CL · 2026-05-20 Cached

SciCustom is a framework for constructing custom scientific benchmarks from large-scale data, enabling fine-grained evaluation of LLMs' scientific capabilities without expert annotation. It uses ontology-grounded knowledge units and voting-based consensus to select relevant benchmarks, demonstrated in chemistry and healthcare.

0 favorites 0 likes
← Back to home

Submit Feedback