Ranking-PE

Ranking-Aware Prompt Optimization for
Multimodal Clinical Diagnosis

1Harvard University2University of California, Santa Cruz3Brown University4Johns Hopkins University5Georgia Institute of Technology

The idea

Align prompt search with ranking quality.

High classification accuracy can conceal poor discrimination when clinical labels are imbalanced. Reflective prompt optimization methods that score individual predictions by correctness can therefore select prompts that perform poorly on ranking metrics.

We introduce Ranking-PE, which aligns three components of reflective prompt optimization with AUROC: a pair-level Pareto score matrix, ranking-shaped feedback, and final prompt selection by validation AUROC. Each matrix row compares a positive case with a negative case, so its average measures pairwise ordering rather than classification accuracy.

On three MIMIC disease tasks, Ranking-PE improves mean test AUROC over Accuracy-PE by 5.8 percentage points on Qwen3-VL-8B with vision-encoder-tuned SFT and 16.2 points on MedGemma-4B. Component ablations examine the contributions of the search objective and feedback.

01 / Motivation

Accuracy and Ranking ask different questions.

The overview connects multimodal clinical inputs, two-stage adaptation, and evaluation: accuracy measures thresholded decisions, while AUROC evaluates positive-negative ordering across thresholds.

Paper overview: chest X-ray and clinical context feed a multimodal model; the evaluation contrasts accuracy with AUROC, followed by visual adaptation and prompt evolution.

02 / Method

From instance correctness to pairwise ordering.

After visual adaptation, Ranking-PE uses the same ranking objective throughout prompt search.

Ranking-PE pipeline: positive-negative pairs form a Pareto score matrix, ranking-shaped feedback guides reflection, and the best prompt is selected by validation AUROC.
01

Pair-level Pareto scores

A row represents a positive-negative pair. Correct ordering scores 1, a tie scores ½, and reversed ordering scores 0. The average over all pairs equals empirical AUROC.

02

Ranking-shaped feedback

The reflector receives the clinical error type, confidence magnitude, and rank context relative to a reference score distribution.

03

AUROC-based selection

The final prompt is selected by validation AUROC. Ranking quality therefore informs both the search and its final decision.

03 / Evaluation

Results across three clinical tasks.

Test performance on Atelectasis, Cardiomegaly, and Consolidation. Compare the prompt optimization recipes within each backbone.

+5.8 AUROC pp

Qwen3-VL-8B + SFT (VE-tuned)

+16.2 AUROC pp

MedGemma-4B

Ranking-PE compared with Accuracy-PE.
Unweighted average across the three disease tasks.

Qwen3-VL-8B + SFT (VE-tuned) · Test AUROC (%)
MethodAtelectasisCardiomegalyConsolidationAvg
No PE63.368.072.668.0
Accuracy-PE63.668.172.468.0
BAcc-Select63.967.671.467.6
Class-Weighted PE64.368.271.568.0
Scalar-AUROC PE66.569.372.969.6
Ranking-PE Ours70.771.379.373.8

04 / Ablation

Each part of the search matters.

On Qwen3-VL-8B + SFT (VE-tuned), changing only final selection to AUROC does not recover the full result in this experiment. Adding pair-level Pareto scores and ranking feedback yields the highest average AUROC among the reported configurations.

Scope of the evidence

These results concern disease tasks from one retrospective cohort. They do not establish performance on other populations or prospective clinical utility. The work studies evaluation and optimization methodology, rather than a clinically validated diagnostic system.

Reference

Citation

@misc{xia2026rankingaware,
  title  = {Ranking-Aware Prompt Optimization for
            Multimodal Clinical Diagnosis},
  author = {Tian Xia and Minghao Liu and Yiqing Liang
            and Laixi Shi and Jiayun Wang},
  year   = {2026},
  eprint = {2609.40361},
  archivePrefix = {arXiv},
  primaryClass = {cs.LG},
  url    = {https://arxiv.org/abs/2609.40361}
}