The limits of accuracy in clinical multimodal adaptation
Multimodal large language models are becoming a fixture in clinical diagnosis, but their training pipelines still treat accuracy as the default objective. In medical data the class distribution can be extreme: a model that always predicts the majority class may clear 90% accuracy while contributing nothing useful to patient care. The paper argues that AUROC offers a better signal—it is threshold-free, ranks positives above negatives, and does not depend on class balance. The authors focus on prompt optimization in multimodal models, showing that reflective methods such as GEPA record per-instance correctness in a scores matrix, and that column averages then drive candidate selection. This accuracy focus can actually degrade ranking performance when the underlying data are imbalanced.
From correctness rows to pairwise ordering
The central contribution is Ranking-PE, which rewrites the scores matrix used by prompt evolution. Instead of one row per evaluation instance with a correctness flag, each row becomes a pairwise ordering over (positive, negative) instance pairs: the entry is 1 if the candidate scores the positive instance higher than its paired negative. The column average of this new matrix is exactly the empirical AUROC, by the Wilcoxon-Mann-Whitney identity. The swap is applied at every layer the prompt evolution search reads from—the scores matrix that decides Pareto dominance, the per-example feedback sent to the reflection language model, and the final candidate selection. No extra model calls are needed and no surrogate loss is introduced.
Ranking-PE in practice
The authors test the approach on three diseases using the MIMIC dataset. When accuracy-based prompt evolution is used, ranking quality can drop; Ranking-PE reverses this trend. On a fine-tuned Qwen3-VL-8B model the method gains +5.8 AUROC points over the accuracy-based recipe. On MedGemma-4B the improvement is larger, +16.2 AUROC points. These gains come without additional inference cost or modified model architecture—the change is purely in the optimization signal.
Visual backbone quality is non-negotiable
Ablation studies reveal that the visual encoder is a make-or-break factor. A medical-grade visual backbone, whether achieved through vision-encoder fine-tuning or medical pretraining, is a prerequisite. Prompt search alone cannot compensate for a weak visual front-end. This finding extends reflective prompt evolution from text-only multimodal settings to the full multimodal clinical decision-making case, where image quality directly affects ranking outcomes.
Takeaways for clinical AI systems
For teams building or fine-tuning multimodal models for medical diagnosis, the paper suggests two concrete steps. First, replace accuracy-based prompt selection with a pairwise-ranking objective such as Ranking-PE; the method adds virtually no computational overhead while directly optimizing the ranking metric that matters in imbalanced clinical settings. Second, invest in a visual encoder that has been tuned for medical imagery—whether through domain-specific fine-tuning or pretraining on clinical datasets—because prompt-level improvements cannot substitute for a capable visual representation. The code and optimization routine are released alongside the paper, providing a ready-to-use swap for existing GEPA-style pipelines.