CLEAR: An Auditable Foundation Model for Radiology Grounded in Clinical Concepts
Paper. CLEAR: an auditable foundation model for radiology grounded in clinical concepts. Nature Biomedical Engineering (2026)
CLEAR asks whether a radiology foundation model can expose clinically named contributions while retaining the flexibility of a large learned representation. The question remains difficult because scale and inspectability pull in different directions. A short concept list can omit findings needed for broad use, while an unrestricted embedding can support many tasks without making its evidence understandable. I read CLEAR as an attempt to make the representation itself available for clinical criticism.
The model was developed using more than 0.87 million chest X-ray–report pairs from 239,391 patients. Its concept space draws on 368,294 report-derived radiological observations. Image–text similarities are combined with text embeddings to construct the final concept-based representation. This is more specific than attaching a list of plausible phrases after prediction. The observations participate in the representation used for downstream tasks, allowing the authors to trace contributions through the constructed pathway.
The distinction between observations and independent clinical factors is important. A large bank of report phrases can contain overlapping descriptions, variations in severity, and related anatomical statements. Its size should not be interpreted as the number of distinct biological variables the model measures. For a reader, the useful question is whether the retrieved or influential observations describe evidence that is present and relevant, and whether their contributions account for the output being inspected.
Reported zero-shot mean AUROC was 78.2% on VinDr compared with 75.0% for CheXzero, and 70.0% on PadChest compared with 66.8%. Linear probing on CheXpert reached a mean AUROC of 87.0%. These comparisons support the feasibility of retaining competitive discrimination with an inspectable representation. Zero-shot evaluation and supervised probing are different conditions, however. A benefit in one should not be presented as evidence that every downstream training arrangement improves.
The most useful experiment concerns enlarged cardiomediastinum. Auditing identified prominent contributions from concepts associated with atelectasis. The intervention retained a subset of mediastinal concepts and refitted the lightweight linear classifier while leaving the image encoder unchanged. AUROC increased from 0.727 to 0.784. This is an example of diagnosis followed by a targeted model revision. It was not simply a clinician correcting one concept value for one patient, and it was not a correction requiring no fitting whatsoever.
That distinction makes the result more usable. The experiment suggests that an explicit concept interface can help formulate and implement a restricted correction without rebuilding the encoder. It does not show that any suspicious concept can be removed safely, or that selecting a clinically relevant subset will always improve transfer. The correction procedure itself becomes part of development. Its performance needs evaluation on data independent of the decisions used to identify and filter the concepts.
The reader assessment adds another kind of evidence: 89.8% of the highest-weighted concepts were judged diagnostically relevant. Relevance is valuable, but it is not identical to presence in the particular image, completeness of the explanation, or necessity for the prediction. A clinically relevant phrase could still be scored incorrectly. A contribution can also be computationally real while representing a correlated finding that is unsuitable for the intended diagnostic claim.
I would therefore separate three questions when examining a CLEAR output. Does the concept score measure the named observation? Does the displayed decomposition accurately describe the selected model output? Does that observation provide appropriate evidence for the clinical task? The architecture can make the second question substantially easier to investigate. It cannot settle the first and third solely through algebra or a readable vocabulary.
The main weakness is the dependence on report-derived semantics. Reports reflect what readers chose to mention, the purpose of the examination, and local documentation conventions. An omitted finding does not necessarily mean absence, and a phrase may contain implications that cannot be verified from the supplied image alone. These are measurement concerns about the concept bank, not reasons to abandon it. They determine which concept outputs require independent annotation and which should retain explicit uncertainty.
A related risk is correlated concept redundancy. If several phrases describe overlapping evidence, individual weights can be difficult to interpret as separate clinical contributions. Filtering one wording may leave closely related information accessible through another. I would want to inspect groups of semantically related observations and the stability of their aggregate role. A large bank creates opportunities for auditing, but it also increases the burden of defining what counts as the same clinical factor.
External evaluation across datasets strengthens the performance evidence, while leaving the destination of a deployment claim to be specified. A model can discriminate well across institutions and still require different probability calibration or operating thresholds. It can also struggle with particular findings despite a strong average. The reported limitation for consolidation is a useful reminder to keep finding-specific performance visible rather than allowing the foundation-model label to imply uniform competence.
This paper supports Auditable-by-Design Medical AI because the concept interface permits a reviewer to inspect a dependency and implement a concrete revision. It also complicates a simple interpretation of Clinical Concepts and Concept-Based Interpretability: a concept-based representation may be broad and editable while its individual values still require semantic validation. Named dimensions are an interface for investigation, not completed evidence of correct clinical reasoning.
The connection to Explanation Faithfulness Versus Plausibility helps distinguish the decomposition, reader assessment, and intervention as separate tests. For my work, I would preserve the encoder, concept bank, text-embedding model, downstream coefficients, and any filtering rule as one versioned pipeline. I would then evaluate concept measurement, computational contributions, and the consequences of a proposed correction independently. That would make an unexpected prediction reproducible and a claimed improvement reviewable.