Dear LISA 2026 Challenge Participants, The Test Phase final ranking methodology for each of the three LISA 2026 Challenge tasks is described below. ### Task 1a – Final Ranking For each image, each performance metric (`Accuracy_weighted`, `F1_weighted`, `F2_weighted`, `Recall_weighted`, `Precision_weighted`) is ordinally ranked across models. For each model, these ordinal ranks of the metrics are used to compute an unweighted geometric rank product for each subject. Each model then has a distribution made up of all the subjects’ unweighted geometric rank products. The median of these distributions is the model’s final score. Rankings are determined by ordering these median scores (lower is better), and breaking ties by comparing the Inter Quartile Range (lower IQR is better). An uncorrected paired Wilcoxon signed‑rank diagnostic test is used to verify that the ranking order aligns with statistical performance trends. Final team rankings were manually verified against the raw metric distributions of each metric. ### Task 1b – Final Ranking For each image, each subject-level image quality metric (`BRISQUE_delta`, `CLIPIQA_enhanced`, `FID_to_reference`, `LPIPS_pair_mean`, `PSNR_to_reference`) is ordinally ranked across models. For each model, these ordinal ranks of the subject-level metrics are used to compute a weighted geometric rank product for each subject. The team-level metric (FRD) is ordinally ranked across the models. A second weighted geometric rank product is calculated using the ordinal ranks of the team-level metric and the subject-level weighted geometric rank products as factors. Each model then has a distribution made up of all the subjects’ weighted geometric rank products which now include all 6 subject and team level metrics. The median of these distributions is the model’s final score. Rankings are determined by ordering these median scores (lower is better), and breaking ties by comparing the Inter Quartile Range (lower IQR is better). An uncorrected paired Wilcoxon signed‑rank diagnostic test is used to verify that the ranking order aligns with statistical performance trends. The weights were derived empirically. Final team rankings were manually verified against the raw metric distributions of each metric. ### Task 2 – Final Ranking For each image, all structure‑specific segmentation metrics (5 metrics: DSC, HD, HD95, ASSD, RVE; 11 brain structures: Thalamus L&R, Caudate L&R, Corpus Callosum, Hippocampus L&R, Lentiform L&R, Ventricle L&R; 55 total metrics) are ordinally ranked across models. For each model, these ordinal ranks of the metrics are used to compute an unweighted geometric rank product for each subject. Each model then has a distribution made up of all the subjects’ unweighted geometric rank products. The median of these distributions is the model’s final score. Rankings are determined by ordering these median scores (lower is better), and breaking ties by comparing the Inter Quartile Range (lower IQR is better). An uncorrected paired Wilcoxon signed‑rank diagnostic test is used to verify that the ranking order aligns with statistical performance trends. Final team rankings were manually verified against the raw metric distributions of each metric. Best regards, The LISA 2026 Challenge Organizers

Created by LISA Challenge LISA_mri_challenge

LISA 2026 Challenge – Test Phase Final Evaluation and Ranking Methodology page is loading…