Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)

Paper. Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV). ICML paper

A pixel attribution can show where a prediction is locally sensitive without identifying the property that matters there. In a medical image, the same region can contain a lesion margin, acoustic texture, and an annotation. TCAV addresses a different question: can a researcher express a hypothesis using examples of a human-named concept, then quantify the model’s sensitivity to that concept? This was an important shift from inspecting a heatmap to testing an explicit semantic hypothesis.

The question remained open because a neural network’s internal coordinates usually lack stable human meanings. A neuron need not correspond to a clinical finding, and a finding can be distributed across many activations. TCAV uses the representation as a space in which a concept direction can be estimated. The diagnostic model remains fixed; the researcher trains an additional classifier to describe a distinction within one of its layers. This makes the procedure an audit of an existing model rather than a new diagnostic architecture.

Concept examples and random counterexamples are passed through the network to obtain their activations at a selected layer. A linear classifier separates the two groups. Its normal vector, oriented toward the concept examples, defines a Concept Activation Vector. The downstream model’s directional derivative along that vector gives a local sensitivity. The TCAV score then summarizes the fraction of examples in a target class for which that sensitivity is positive. The paper repeats this process with random sets and applies statistical testing.

Writing the quantities explicitly helps prevent overinterpretation. If \(a(x)\) is the selected activation, \(h_k(a)\) the target class logit, and \(v_C\) the concept direction, the local quantity is \(\nabla_a h_k(a(x))^\top v_C\). Its sign asks whether a small movement in that direction increases the selected logit. The class-level score counts positive signs. It is not the average magnitude of the derivative, a fraction of explained variance, or the proportion of predictions caused by the concept.

The original applications include natural-image classification and a diabetic-retinopathy model. They demonstrate that concept-based tests can expose class-level sensitivities that are difficult to express through individual pixels. That is a methodological demonstration with considerable clinical relevance. It does not establish that every plausible clinical concept has a reliable linear direction, or that the tested concept vocabulary exhausts the information used by a model. Method and experiments.

The important control is that the audited classifier is held fixed while concept directions are fitted. Otherwise, a change in a TCAV score could mix changes in the diagnostic system with changes in the measurement procedure. Even with a fixed system, the layer, concept examples, comparison set, and selected output define the experiment. Reporting “TCAV for irregular margins” without those details leaves the actual measurement underspecified.

A careful reader should first question the meaning of the fitted direction. Suppose concept-positive ultrasound images mostly come from one scanner and the controls from another. A highly accurate separator could represent scanner differences, margin differences, or both. Calling the direction “irregular margin” would not resolve that ambiguity. I would therefore inspect concept examples as a dataset in their own right, document their provenance, and evaluate whether separation persists on patients and acquisition settings excluded from concept fitting.

The random counterexamples are also substantive controls. Different counterexamples can change which visual distinctions make the concept separable. Repeated fitting helps reveal instability, but repeated random sets cannot correct a confound shared by every concept example. Statistical significance answers a question about the specified repeated procedure. It does not certify the clinical meaning of the direction, and testing many concepts or layers introduces additional opportunities to select an appealing result.

A second weakness concerns the move from local sensitivity to clinical reliance. A small activation-space movement need not correspond to a realizable image with only one finding changed. The direction can intersect several correlated properties or move away from representations produced by real images. A positive derivative therefore establishes a computational response under the specified direction. It does not establish that the clinical finding is necessary, sufficient, or causally responsible for the patient’s disease.

The distinction between sign and magnitude matters practically. A score near one can arise when nearly every derivative is positive but small. A lower score can conceal large effects in a clinically important subset. Class-level aggregation can also obscure opposing behavior across scanners or disease presentations. I would inspect the distribution of local sensitivities and relevant subgroups before treating a single TCAV score as a complete description of model behavior.

For my work, a useful first application would compare a clearly defined clinical finding with a suspected documentation cue. The concept sets would need independent annotations, patient separation, and acquisition information. I would assess how results change across reasonable layers and control sets, then examine whether the predicted direction of response agrees with carefully validated image or internal interventions. This is a proposed audit sequence, not an intervention guarantee supplied by TCAV.

Representation-Level Auditing places TCAV between simple probing and direct behavioral tests. A probe asks whether information is recoverable; TCAV additionally asks whether the downstream output is locally sensitive along a fitted direction. The method supports the note’s distinction between accessibility and sensitivity, while leaving the clinical interpretation of that sensitivity to be validated.

Clinical Concepts and Concept-Based Interpretability explains why a named variable needs an observation protocol. Intervention-Based Auditing supplies the next challenge: even a successful edit must be checked for collateral changes. TCAV is most useful to me as a disciplined way to formulate and prioritize hypotheses. Its value is greatest when a concept score leads to a more discriminating experiment, rather than ending the inquiry.