Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models Conference

Yang, X, Zhong, Z, Collins, S et al. (2026). Reliability Stress Tests and Decision-Time Routing for Chest X-ray Vision-Language Models . 428-433. 10.1109/CHASE69719.2026.00073

cited authors

  • Yang, X; Zhong, Z; Collins, S; Baird, G; Wang, X; Jiao, Z

authors

abstract

  • Medical vision-language model (VLM) evaluation is sensitive to workflow design, prompting strategy, and benchmark construction, yet most studies treat these factors in isolation. We introduce a reliability stress test for chest X-ray interpretation built on two balanced datasets (a private report-backed set and a curated MIMIC subset). Three medical VLMs - CheXagent, MedGemma-4B, and MedGemma-27B - are evaluated across three prompt styles and two workflows (single-VLM and multi-agent), producing 36 configurations. We show that exact-match accuracy alone can overstate the effectiveness of conservative models that default to "Normal"predictions. Diagnostic reliability also depends heavily on model family and scale: multi-agent reasoning helps some configurations but hurts others. Building on these observations, we propose a decision-time routing framework that selectively escalates to multi-agent inference only when beneficial, improving the cost-quality trade-off over fixed workflows. Our results highlight the need for evaluation protocols that jointly consider prompt sensitivity, failure-mode diversity, and workflow choice before clinical deployment.

publication date

  • January 1, 2026

Digital Object Identifier (DOI)

start page

  • 428

end page

  • 433