Quote of the day

A person who never made a mistake never tried anything new.

- Albert Einstein

Contact

Guideline-Grounded Clinical Language Models with Calibrated Abstention for Primary Care Triage

Open Medicine & Healthcare

Guideline-Grounded Clinical Language Models with Calibrated Abstention for Primary Care Triage

Guideline-Grounded Clinical Language Models with Calibrated Abstention for Primary Care Triage

Description

Design and evaluate a primary-care triage assistant that answers only when the relevant clinical recommendation can be retrieved and that explicitly quantifies uncertainty.

Details

Context and Problem Statement

Medical LLMs may score well on exam questions while failing in real settings where information is incomplete and error costs are asymmetric. In primary care, the objective is safe orientation and reliable red-flag detection rather than autonomous diagnosis.

Research Question

How can calibrated abstention control high-severity triage errors, and how should clinical safety be measured beyond average accuracy?

Proposed Approach

Use RAG over validated guidelines with mandatory citation, explicit red-flag rules, uncertainty estimation through sampling consistency, and conformal prediction. Evaluate scenarios stratified by severity and report low-harm and high-harm errors separately.

Expected Contribution

A safety-oriented evaluation methodology and evidence on the benefit of calibrated abstention compared with increasing model size.

Expected Prototype

A clinician-facing triage assistant showing proposed orientation, evidence, detected red flags, confidence, and mandatory escalation when appropriate.

Datasets

HealthBench, public clinical benchmarks, WHO and professional-society guidelines, and physician-validated synthetic cases.

Challenges

Uncertain ground truth, misuse outside professional oversight, bias, medico-legal responsibility, and software-medical-device regulation.

Research Question

Does calibrated abstention reduce high-harm triage errors more effectively than increasing model capacity under a comparable computational budget?

Innovation

The project replaces global accuracy with severity-aware safety evaluation and statistical coverage guarantees.

Expected Deliverable

A supervised clinical-triage prototype, a safety-oriented benchmark, and a clinician-reviewed failure-mode analysis.

Technologies

Large Language Models Retrieval-Augmented Generation Conformal Prediction and Uncertainty Estimation Structured Clinical Evaluation Trustworthy AI

Required Skills

  • NLP and RAG Engineering
  • Statistics and Uncertainty Quantification
  • Foundational Biomedical Knowledge
  • Expert-Based Evaluation Study Design

Datasets

  • HealthBench (OpenAI, open evaluation benchmark)
  • WHO and professional-society clinical guidelines
  • Synthetic clinical cases validated by physicians

Morocco & Africa Relevance

Low physician density in many Moroccan and African regions increases interest in triage support, making safe abstention essential.