VLMs can score well on benchmarks, while silently erasing meaningful terms and including hallucinate bias [P]

While working with VLMs for report generation on chest x-rays (RRG), we noticed that evaluation metrics are flawed. Flawed in a sense where they rewarded repetitive templates, reports without clinical terms and reports which were "normal" with high scores on benchmark metrics. Also, clinically meaningful but rare words were erased leaving the generated report looking repetitive and boring. Importantly, of no clinical utility. In the paper below, we discuss this behaviour of VLMs for RRG and introduce a framework to actually measure the erasure of terms and introduction of biased terms. Paper: Measuring What VLMs Don't Say: Validation Metrics Hide Clinical Terminology Erasure in Radiology Report Generation Link: Reference Paper Url: arxiv.org/abs/2603.01625

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论