Guideline-Grounded Clinical Language Models with Calibrated Abstention for Primary Care Triage
Guideline-Grounded Clinical Language Models with Calibrated Abstention for Primary Care Triage
Description
Details
Context and Problem Statement
Medical LLMs may score well on exam questions while failing in real settings where information is incomplete and error costs are asymmetric. In primary care, the objective is safe orientation and reliable red-flag detection rather than autonomous diagnosis.
Research Question
How can calibrated abstention control high-severity triage errors, and how should clinical safety be measured beyond average accuracy?
Proposed Approach
Use RAG over validated guidelines with mandatory citation, explicit red-flag rules, uncertainty estimation through sampling consistency, and conformal prediction. Evaluate scenarios stratified by severity and report low-harm and high-harm errors separately.
Expected Contribution
A safety-oriented evaluation methodology and evidence on the benefit of calibrated abstention compared with increasing model size.
Expected Prototype
A clinician-facing triage assistant showing proposed orientation, evidence, detected red flags, confidence, and mandatory escalation when appropriate.
Datasets
HealthBench, public clinical benchmarks, WHO and professional-society guidelines, and physician-validated synthetic cases.
Challenges
Uncertain ground truth, misuse outside professional oversight, bias, medico-legal responsibility, and software-medical-device regulation.
Research Question
Innovation
Expected Deliverable
Technologies
Required Skills
- NLP and RAG Engineering
- Statistics and Uncertainty Quantification
- Foundational Biomedical Knowledge
- Expert-Based Evaluation Study Design
Datasets
- HealthBench (OpenAI, open evaluation benchmark)
- WHO and professional-society clinical guidelines
- Synthetic clinical cases validated by physicians