Weight Space Correlation Analysis: Quantifying Feature Utilization in Deep Learning Models

Paper. Weight Space Correlation Analysis: Quantifying Feature Utilization in Deep Learning Models. MIDL 2026 paper

A representation can reveal the scanner that produced an image without the diagnostic head using scanner identity. This distinction is easy to state and difficult to investigate. A successful metadata probe establishes that information is accessible under the probe’s learning procedure. It does not establish what the original model does with that information. I read this paper because it introduces a practical intermediate test between metadata decoding and an intervention on the input or model.

The question remained open because both common alternatives have limitations. Probing can produce an alarming result even when the task head ignores the decoded information. Counterfactual editing can test prediction changes more directly, but acquisition factors such as scanner characteristics may be distributed throughout an image and difficult to manipulate selectively. Weight Space Correlation, or WSC, asks whether the clinical and metadata prediction heads use similar directions in a shared representation.

The method assumes a backbone followed by linear classification heads. Auxiliary heads predict metadata from frozen embeddings. Principal component analysis is applied to the training embeddings, and the clinical and auxiliary weight vectors are projected into that common subspace before cosine similarity is computed. The principal components therefore describe variation in the representations, not a principal-component summary fitted independently to each head. This detail matters because the comparison relies on a shared coordinate system.

The experiments include fetal-plane classification, deliberately induced scanner-related shortcuts, and analysis of a spontaneous-preterm-birth model using cervical ultrasound. Controlled shortcut settings provide a useful check because the analyst knows how the association was introduced. The clinical application then examines whether geometric alignment favors clinically relevant or acquisition-related factors. The reported pattern includes alignment with cervical length and weak alignment with scanner identity. These are global observations about the examined representation and heads.

One particularly instructive qualification concerns birth weight. In the published paper, a high WSC value for that probe is discounted because the probe itself predicts birth weight poorly. This prevents a geometric quantity from being interpreted without checking whether the associated direction represents the named variable. A weight vector exists even for an unsuccessful predictor. Its alignment can be numerically striking while conveying little useful information about the intended clinical concept.

What convinces me is this separation of accessibility and alignment. A useful audit should report both. If a nuisance variable is decoded well but aligns weakly with the task head, the result weakens the simple argument that recoverability proves reliance. If decoding and alignment are both strong, the audit has identified a more focused hypothesis. Neither outcome should be compressed into an unconditional declaration that the model is trustworthy or that a shortcut has been causally established.

Cosine similarity also needs a precise interpretation. It measures an angle in the chosen coordinates, not the correlation of model outputs and not the effect of changing a clinical variable. As a constructed example, suppose independent representation coordinates have variances nine and one. Scores proportional to their sum and difference have orthogonal weight vectors, yet their covariance is eight. Unequal variation in the representation can therefore make weight geometry and score association answer different questions.

That example does not invalidate the proposed statistic. It identifies what I would want to examine before treating its magnitude as a direct amount of feature use: representation scaling, covariance structure, retained components, and the behavior of the actual outputs. The PCA step makes the analysis more data-aware, but the choice of subspace still matters. A direction with little overall variance might be important for a rare subgroup, and a global projection can make such dependence harder to detect.

The clinical meaning of the probe is another limitation. A scanner head can use anatomy, patient mixture, or acquisition settings correlated with scanner identity. Alignment with that head then identifies overlap with a predictive contrast, rather than isolating a single physical scanner mechanism. Conversely, a weak linear probe can miss information expressed through interactions. These are reasons to evaluate probe performance across relevant groups and to avoid interpreting a label attached to a head as a complete semantic description of its weights.

Low WSC should therefore be treated as limited negative evidence. It does not rule out every form of shortcut reliance, nor does it guarantee transfer to another scanner. A new device may change the visibility or rendering of valid diagnostic features even when scanner identity was not an explicit decision cue. Distinguishing these failure mechanisms is useful, but the weight comparison alone cannot conclusively assign the cause of a later performance drop.

For an ultrasound audit, I would preserve the original task predictor, fit probes on a separate development procedure, and document the layer, pooling, normalization, and projection choices. I would compare alignment across seeds and clinically relevant subgroups, including a null or shuffled-label probe. A promising dependency would then motivate a matched acquisition comparison or a carefully validated intervention. These follow-up tests would assess whether changing the suspected factor changes the original prediction under interpretable conditions.

The paper directly supports Representation-Level Auditing, which distinguishes accessible information, representation structure, and sensitivity reaching the final output. It also complicates any overly literal reading of Geometry of Representation Spaces: an angle is meaningful only relative to coordinates, data variation, and the function attached to the representation. Geometric similarity does not supply clinical semantics by itself.

The connection to Intervention-Based Auditing explains where I would place WSC in a research workflow. It can prioritize dependencies worth testing when direct edits are expensive or difficult. Its value is strongest as a disciplined screening measurement with explicit prerequisites. For my work, the desired output would be a small set of well-supported reliance hypotheses, accompanied by probe validity and uncertainty, that can be challenged using the fixed diagnostic model.