Tag
This study develops an ambiguity taxonomy to evaluate large language model performance on clinical registry abstraction from unprocessed EMR data, finding that LLM accuracy is significantly lower than human abstractors and declines as task ambiguity increases.