Tag
This paper introduces a graph-grounded harness for vision-language models to improve accuracy in answering topology questions about Piping and Instrumentation Diagrams by recovering evidence graphs from images.