The dataset is min–max normalised, but the maximum voxel value can be highly sensitive to noise and outliers. This may cause arbitrary differences in overall image brightness, even between scans of the same modality and field strength. In comparison, P95 or z-score normalisation produces more consistent tissue contrast across scans. Our main concern is that the current evaluation may place too much weight on predicting this unstable maximum value. For example, even if a model correctly predicts most voxels in a stable z-score or P95-normalised space, a small error in the estimated maximum will rescale the entire image when converted back to the min–max scale. This increases the error across nearly all voxels. As a result, the evaluation may reflect how accurately the model predicts a single extreme voxel value, rather than how well it reconstructs the underlying anatomy and tissue contrast.

Created by YULIANG HUANG hyliang
Thank you very much for raising this important point. As the evaluation protocol and challenge rules were fixed before the test phase, we are unfortunately unable to modify the normalization or evaluation procedure at this stage, in order to ensure fairness and consistency for all participating teams. We have also noticed this issue during the challenge and agree that intensity normalization may influence similarity-based metrics. In our final evaluation and analysis, we will therefore pay particular attention to the consistency between image-similarity metrics and anatomy/segmentation-based evaluation results, rather than interpreting a single metric in isolation. Your comment is very helpful, and we will consider more robust normalization and evaluation strategies in future editions of the challenge. Thank you again for your thoughtful feedback and contribution to improving MRIxFields.

Concern about evaluation bias from min–max normalisation page is loading…