Classifications in Brief: Kellgren-Lawrence Classification of Osteoarthritis
Paper. Classifications in Brief: Kellgren-Lawrence Classification of Osteoarthritis. Clinical Orthopaedics and Related Research (2016)
KL grade is easy to treat as a natural property of an image because it appears so often as a target in knee-AI studies. This review makes it harder to ignore that the grade is a constructed summary. Its meaning depends on a combination of radiographic findings, a grading convention, and a reader’s application of that convention. Before asking a model to predict KL, I need to understand what agreement with that label actually establishes.
The authors review the classification’s history, intended purpose, reliability evidence, and limitations. They do not introduce a new image classifier or conduct one unified comparison on a single modern dataset. Numerical reliability estimates come from studies with different populations, views, and reader arrangements. That makes this paper useful for understanding the range of conditions under which KL has been applied, while limiting any attempt to quote one reliability coefficient as an intrinsic property of the scale.
KL organizes radiographic osteoarthritis into five ordered categories using findings including osteophytes, joint-space narrowing, sclerosis, and bony deformity. The distinction between early grades gives an important role to whether osteophytes are definite, while higher grades combine more marked narrowing and additional structural changes. It is a composite convention, not a direct measurement of cartilage thickness or a count of independent abnormalities. Classification and validation review.
The setup matters because an ordinal summary can make a disease appear more one-dimensional than its component observations. A knee with substantial narrowing and few definite osteophytes does not fit comfortably into a progression story that privileges osteophyte formation. The review discusses this difficulty explicitly. For AI, this means that an unusual feature combination may be clinically meaningful even if it is awkward to encode in the target label.
The reliability evidence reinforces the dependence on protocol. The review cites an investigation reporting interobserver ICCs of about 0.54 for Rosenberg views and 0.38 for AP views. Other cited studies reported higher agreement. These values should remain attached to their study conditions; selecting only the lowest values would unfairly portray KL as uniformly unreliable. Their practical lesson is that the imaging view, case mixture, and reader procedure can change apparent agreement substantially.
A model evaluated against one reader therefore answers a different question from a model evaluated against an adjudicated consensus. Agreement with a reader may partly reflect that reader’s thresholds. Consensus can reduce some individual variation while retaining shared assumptions or systematic errors. Neither reference automatically reveals an error-free underlying grade. When reporting performance, I would specify who produced the target, what information they saw, and how disagreements were handled.
It is also too strong to say that interreader agreement gives a universal numerical ceiling on model performance. A model might reproduce a consensus more consistently than an individual reader does, or learn systematic features of the labeling process. Conversely, very high agreement with noisy labels can be misleading if the noise is predictable from acquisition or annotation conventions. Reader variability provides context for interpretation, not a simple upper bound that can be transferred between datasets.
Ordinal evaluation needs more detail than overall accuracy. Confusing adjacent grades and confusing the extremes are different patterns, but the clinical meaning of an error also depends on which decision boundary it crosses. Weighted agreement measures encode a choice about how to penalize errors. Squared error on grade numbers assumes a particular spacing between categories. Those choices can be reasonable, but they should be described as analytical decisions rather than properties established by the KL definitions.
Longitudinal use introduces another limitation. A change in a coarse grade can miss smaller structural changes within a category, while a borderline case can cross a category boundary with modest differences in appearance or interpretation. Predicting future KL change is therefore prediction of a reference-defined radiographic endpoint. It is not automatically prediction of symptom progression, cartilage loss measured continuously, or future treatment need. That distinction is essential when interpreting the multi-joint-space-width paper elsewhere in this corpus.
A careful reader would also ask whether replacing KL solves the problem or merely changes it. Component scores preserve information about individual findings, but require more annotations and retain uncertainty. A continuous measurement avoids some category boundaries while introducing questions about positioning, calibration, and measurement location. The review does not establish a universally superior replacement. The appropriate target depends on what the study needs to measure and what information the acquisition supports.
For a new KL model, I would keep the complete confusion matrix and report performance at clinically relevant boundaries. I would examine cases where component findings and the assigned grade appear discordant. Those cases should receive a structured review, not automatic exclusion as “bad labels.” A model that follows visible narrowing while disagreeing with an osteophyte-centered convention may be measuring something useful, but a separate reference is needed before claiming that it is more correct.
The OARSI atlas provides a natural companion because it makes component observations explicit. Clinical Feature Annotation and Multi-Task Learning explains how those observations can become supervision without being conflated with diagnosis or action. Together, they suggest preserving the component evidence alongside a global grade, so that a model error can be localized to a finding or to the rule combining findings.
Label Quality and Interobserver Variability is the central study-note connection. This review supplies clinical examples of why agreement depends on definitions, information, and readers. Clinical Validity and Clinical Utility marks the next boundary: accurate reproduction of radiographic grading does not establish better patient outcomes. I take KL as a useful, established measurement convention whose limitations must remain visible in the model’s claims.