Knee-xRAI: An Explainable AI Framework for Automatic Kellgren-Lawrence Grading of Knee Osteoarthritis
Paper. Knee-xRAI: An Explainable AI Framework for Automatic Kellgren-Lawrence Grading of Knee Osteoarthritis. arXiv manuscript
Knee-xRAI asks whether automatic Kellgren-Lawrence grading can expose the radiographic findings behind a grade. This is an appealing question because a single ordinal output compresses several observations into one category. A disagreement becomes easier to investigate if the system also provides joint-space measurements, osteophyte assessments, and sclerosis outputs. The clinical stakes motivate careful evaluation, but the paper does not establish a treatment policy. A radiographic grade alone should not be interpreted here as determining surgery.
The framework combines a U-Net++ joint-space module, a network for site-specific osteophyte grading, and a texture-based sclerosis component. Their outputs form a 50-dimensional structured vector. That vector feeds two different classifiers: an XGBoost path using structured features alone, and a ConvNeXt hybrid that also receives image features. This two-path design is central to the interpretation. The model with the clearest feature-level account is not the same computational object as the model with the strongest reported grading performance.
Evaluation used 8,260 OAI-derived radiographs, with a held-out test split of 1,656 images. Feature annotations were available on smaller subsets. The structured XGBoost path reached test quadratic weighted kappa of 0.6294, while the hybrid reached 0.8436 and an AUC of 0.9017. Those figures establish a substantial difference between the evaluated paths. They do not, by themselves, quantify an unavoidable cost of transparency, because architecture, available information, and optimization differ together.
The paper also tests the hybrid directly. At inference, zeroing the structured vector reduced kappa by approximately 0.098, and permuting it across images reduced kappa by approximately 0.284. These experiments leave the trained hybrid fixed and alter its structured input. They provide evidence that the hybrid uses that pathway and that correspondence between the features and image matters. This is stronger evidence about the hybrid than displaying SHAP values from the separate XGBoost classifier.
The intervention nevertheless has a specific meaning. Zeroing standardized features does not necessarily represent the clinical absence of joint-space narrowing, osteophytes, or sclerosis. It can represent a reference numerical value. Permuting features creates image–feature combinations that may be implausible. A performance decline therefore demonstrates sensitivity to those input manipulations, but it does not measure the effect of selectively changing a radiographic finding while preserving everything else.
This matters when interpreting the dominance of joint-space information. A strong response to disrupting that feature family supports its computational importance under the test. It does not establish that every individual joint-space measurement is correct or that the model uses it in the same way a clinician would. It also does not establish that the image encoder contains no overlapping information. The two paths can carry redundant, complementary, or conflicting evidence.
SHAP explanations need the same model-specific discipline. An attribution for the XGBoost path describes that model under its explanation assumptions. It cannot be transferred to the hybrid merely because both receive the structured vector. The hybrid can obtain information directly from the image, including information not named in the structured outputs. A clinician-facing display should make clear which path produced the prediction and which calculation produced its explanation.
The feature vector itself also needs interpretation. Fifty numerical variables are not automatically fifty independently validated clinical concepts. A measurement, an ordinal grade, a texture statistic, and a learned score have different semantics. Their units and uncertainty should be visible. In particular, the paper notes that the image format lacks calibrated DICOM pixel spacing, so joint-space measurements are pixel-based unless an external scale is supplied. That limits any direct interpretation in physical units.
The most serious empirical weakness is uneven evidence quality across modules. The feature-label subsets are small relative to the grading dataset, inter-annotator reliability was not formally quantified, and some component performance is modest. A high final grading score can coexist with an unreliable intermediate output if other inputs compensate for it. That makes component validation essential for an interface whose value depends on readers trusting those intermediate measurements.
I would also want patient-level split provenance before making a strong independence claim. The reviewed methods describe radiograph partitions, but the reported counts alone do not establish whether every related image is separated appropriately. This is an unresolved reporting question, not proof of leakage. Likewise, the absence of external validation leaves institutional and acquisition transfer open. Motivation from a resource-constrained setting does not itself demonstrate performance in that setting.
A useful follow-up would test three things separately: measurement agreement, grading performance, and the usefulness of the displayed evidence. For measurements, I would examine errors by severity, acquisition quality, and anatomical site. For the hybrid, I would add clinically constrained feature edits and check whether the supplied image still supports the altered feature combination. For the interface, readers would need to identify erroneous outputs and make decisions with and without assistance, rather than merely rate explanations as understandable.
The paper supports Auditable-by-Design Medical AI by exposing intermediate outputs that can be questioned. It also makes the note’s requirement to preserve the complete prediction pathway concrete: an audit of the structured path cannot stand in for an audit of the hybrid. The actual model version, feature extraction, and selected path must remain linked to each displayed score.
Label Quality and Interobserver Variability explains why component annotations are part of the measurement apparatus, while Intervention-Based Auditing explains the limits of zeroing and permutation. For my gallbladder work, the transferable design is a system with independently testable findings and interventions on the deployed predictor. The remaining obligation is to show that the named findings mean what they claim to mean and that the explanation covers the pathway actually used.