Pair-level Pareto scores
A row represents a positive-negative pair. Correct ordering scores 1, a tie scores ½, and reversed ordering scores 0. The average over all pairs equals empirical AUROC.
Ranking-PE
The idea
High classification accuracy can conceal poor discrimination when clinical labels are imbalanced. Reflective prompt optimization methods that score individual predictions by correctness can therefore select prompts that perform poorly on ranking metrics.
We introduce Ranking-PE, which aligns three components of reflective prompt optimization with AUROC: a pair-level Pareto score matrix, ranking-shaped feedback, and final prompt selection by validation AUROC. Each matrix row compares a positive case with a negative case, so its average measures pairwise ordering rather than classification accuracy.
On three MIMIC disease tasks, Ranking-PE improves mean test AUROC over Accuracy-PE by 5.8 percentage points on Qwen3-VL-8B with vision-encoder-tuned SFT and 16.2 points on MedGemma-4B. Component ablations examine the contributions of the search objective and feedback.
01 / Motivation
The overview connects multimodal clinical inputs, two-stage adaptation, and evaluation: accuracy measures thresholded decisions, while AUROC evaluates positive-negative ordering across thresholds.

02 / Method
After visual adaptation, Ranking-PE uses the same ranking objective throughout prompt search.

A row represents a positive-negative pair. Correct ordering scores 1, a tie scores ½, and reversed ordering scores 0. The average over all pairs equals empirical AUROC.
The reflector receives the clinical error type, confidence magnitude, and rank context relative to a reference score distribution.
The final prompt is selected by validation AUROC. Ranking quality therefore informs both the search and its final decision.
03 / Evaluation
Test performance on Atelectasis, Cardiomegaly, and Consolidation. Compare the prompt optimization recipes within each backbone.
Qwen3-VL-8B + SFT (VE-tuned)
MedGemma-4B
Ranking-PE compared with Accuracy-PE.
Unweighted average across the three disease tasks.
| Method | Atelectasis | Cardiomegaly | Consolidation | Avg |
|---|---|---|---|---|
| No PE | 63.3 | 68.0 | 72.6 | 68.0 |
| Accuracy-PE | 63.6 | 68.1 | 72.4 | 68.0 |
| BAcc-Select | 63.9 | 67.6 | 71.4 | 67.6 |
| Class-Weighted PE | 64.3 | 68.2 | 71.5 | 68.0 |
| Scalar-AUROC PE | 66.5 | 69.3 | 72.9 | 69.6 |
| Ranking-PE Ours | 70.7 | 71.3 | 79.3 | 73.8 |
04 / Ablation
On Qwen3-VL-8B + SFT (VE-tuned), changing only final selection to AUROC does not recover the full result in this experiment. Adding pair-level Pareto scores and ranking feedback yields the highest average AUROC among the reported configurations.
These results concern disease tasks from one retrospective cohort. They do not establish performance on other populations or prospective clinical utility. The work studies evaluation and optimization methodology, rather than a clinically validated diagnostic system.
Reference
@misc{xia2026rankingaware,
title = {Ranking-Aware Prompt Optimization for
Multimodal Clinical Diagnosis},
author = {Tian Xia and Minghao Liu and Yiqing Liang
and Laixi Shi and Jiayun Wang},
year = {2026},
eprint = {2609.40361},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2609.40361}
}