FirsthandHealth
medRxiv PreprintsInternational6 October 2026

Phantom Fairness: A Reproducible Audit of How Automatically Extracted Labels Can Conceal Demographic Disparities in Chest Radiograph Classifiers

This is an official announcement record

Firsthand records what medRxiv Preprints announced and links to the original. The wording below is theirs, not ours.

Fairness audits of chest radiograph artificial intelligence usually score subgroup performance against labels extracted from radiology reports by natural language processing (NLP). We tested whether these labels can make a model look fairer than it is, a failure we call phantom fairness. On 4,376 NIH ChestX-ray14 images with both NLP and radiologist-adjudicated labels, we audited classifiers from three backbones by sex, age, projection and older women. NLP labels missed 54% to 59% of radiologist-confirmed airspace opacity, pneumothorax and nodule or mass, and every fracture. These misses were
— medRxiv Preprints
Read the official announcement

Opens www.medrxiv.org

More from medRxiv Preprints

This content is for informational purposes only and is not medical advice. It is not intended to diagnose, treat, cure, or prevent any disease. Consult a healthcare professional before starting any supplement, treatment, or program — especially if you are pregnant, nursing, taking medication, or managing a health condition.