Underdiagnosis Bias of Artificial Intelligence Algorithms Applied to Chest Radiographs in Under-Served Patient Populations
Paper. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations.
The question is whether a chest-radiograph classifier makes a particular consequential error more often for some patient groups: predicting no finding when the dataset reference indicates an abnormality. This is more specific than asking whether overall accuracy differs between groups. A system used to deprioritize apparently normal examinations could distribute missed opportunities for review unevenly even while achieving a respectable aggregate score. The paper therefore begins with an error direction that matters to a possible clinical workflow.
That question remained open because strong benchmark performance does not describe how errors are distributed. An overall metric combines patients with different disease presentations, acquisition conditions, and frequencies of abnormal findings. It can improve while a smaller group receives little benefit. Moreover, separate analyses of sex, age, or race can conceal intersections where the pattern differs from either parent group. A fairness evaluation needs to identify both its denominator and the consequences attached to its errors.
The authors examine models trained using MIMIC-CXR, CheXpert, the NIH chest-radiograph dataset, and pooled data. They compare underdiagnosis across available demographic and insurance attributes, including intersections, and repeat training over five random seeds. Their central endpoint is the false-positive rate for the no-finding label. The terminology needs care: a false positive for “no finding” corresponds to falsely reassuring an examination that the reference labels as abnormal. It is not the same endpoint as a false positive for pneumonia.
Within a subgroup, the relevant denominator is examinations with a reference finding, rather than all examinations or all model errors. The reported patterns include higher underdiagnosis among several underserved groups, although the affected groups and directions are not identical in every dataset. In particular, the NIH findings should not be compressed into a claim that the same demographic ordering holds everywhere. Intersectional results motivate looking beyond single attributes, but they do not establish that every smaller intersection necessarily experiences a larger disparity.
The useful comparison holds the trained predictor fixed while examining its behavior across groups. However, the groups themselves are not matched experiments. Disease composition, severity, image quality, and documentation can differ together. A higher subgroup error rate establishes a performance disparity under the evaluated reference and operating rule. It does not, by itself, identify whether the mechanism is underrepresentation, acquisition differences, label quality, disease spectrum, or some combination. These possibilities call for different repairs.
The broad no-finding target is both the paper’s strength and an important limitation. It resembles a potentially consequential triage decision, but it combines abnormalities with different visibility and urgency. Missing a subtle chronic finding and missing an acute abnormality can enter the same numerical endpoint while having different clinical implications. I would therefore read the result as a reason to investigate missed findings in detail, rather than as a complete measurement of clinical underdiagnosis.
The reference standard also bounds the claim. Report-derived labels measure what was documented and then extracted from reports. They are not an independent observation of every patient’s underlying disease. If documentation or extraction quality differs between groups, measured disparities can reflect both model behavior and reference error. This does not make the observed differences unimportant. It changes the next experiment: a stratified review with an independently specified reference could help determine which disagreements represent missed visible disease and which arise elsewhere in the labeling process.
Thresholds deserve equal attention. Underdiagnosis is measured after converting a score into a decision. A model can have useful ranking performance and still produce an unacceptable error rate at the chosen threshold. Moving that threshold generally changes other errors and the number of examinations requiring review. Consequently, a proposed fairness correction should report the resulting sensitivity, specificity, workload, and subgroup uncertainty. Equalizing one rate is not enough to establish that the revised pathway is clinically preferable.
The repeated training runs help show that the conclusion is not simply a story about one checkpoint. They do not resolve every source of uncertainty. Variation across seeds, sampling uncertainty among patients, and uncertainty in the reference labels are different quantities. Sparse intersections may remain poorly characterized even when training behavior is stable. A nonsignificant subgroup difference should therefore not be described as evidence that the groups perform equivalently.
This paper gives a concrete application of Robustness, Subgroup Performance, and External Validation. That note emphasizes conditional performance and the weights connecting subgroup results to an aggregate metric. The study supports its argument that generalization claims need a named population and endpoint. It also complicates any assumption that collecting data from several sources automatically removes disparities: pooled development data still need subgroup evaluation, and they do not substitute for testing the intended destination.
It also belongs beside Dataset Design, Ground Truth, and Reference Standards and Error and Clinical Failure Analysis. The former explains why reference provenance belongs inside the performance claim. The latter separates potential severity from observed harm. These retrospective classification results identify a plausible route to unequal care, but they do not directly measure treatment delay, denied care, or patient outcomes after deployment.
For my own ultrasound work, I would borrow the paper’s choice to start from a directional failure. For example, the relevant question might concern examinations incorrectly cleared of a suspicious finding, with separate accounting for incomplete visualization. I would preserve patient-level denominators, review errors alongside correctly classified controls, and examine presentation and acquisition factors within demographic groups. The eventual claim should specify what was missed, against which reference, at which decision rule, and in which patients. This paper makes that evaluation necessary; it does not supply a universal explanation or correction for the disparities it reveals.