Proposed International Clinical Diabetic Retinopathy and Diabetic Macular Edema Disease Severity Scales

Paper. Proposed International Clinical Diabetic Retinopathy and Diabetic Macular Edema Disease Severity Scales. Ophthalmology (2003)

This paper asks how to make disease severity communicable across clinical settings. That is a different starting point from asking how accurately an algorithm can assign a retinal photograph to a class. I read it because the meaning of the class must be established before agreement with it can be interpreted. A model trained on diabetic retinopathy grades inherits decisions about what observations count, which distinctions matter, and which distinctions the label deliberately compresses.

The practical problem was that detailed research grading systems were difficult to apply and communicate in routine care. The authors wanted a common vocabulary usable by clinicians with different training and equipment. Simplification therefore had a purpose: preserve clinically consequential distinctions while making classification feasible outside specialized research grading. The resulting categories should be understood as an evidence-informed clinical convention. Their usefulness does not depend on pretending that disease naturally divides into five perfectly separated groups.

The development process involved 31 participants from 16 countries. An initial proposal drew on the Early Treatment Diabetic Retinopathy Study and Wisconsin Epidemiologic Study of Diabetic Retinopathy. Participants reviewed proposals, provided ratings through a modified Delphi process, discussed them at a workshop, and reconsidered the revised classifications. The primary result was agreement on the classification scheme. This was not a diagnostic-model experiment with a training set, a held-out test set, and a measured classification accuracy.

The final retinopathy scale has five stages: no apparent retinopathy, mild, moderate, and severe nonproliferative retinopathy, and proliferative retinopathy. Macular edema is classified separately, with further distinctions depending on its location and the examiner’s ability to assess that location. This separate treatment matters for AI task definition. Success at predicting a retinopathy category does not establish that a system has assessed the macula adequately or answered every question relevant to the examination.

The deliberations also expose the construction of a boundary. The paper describes disagreement over whether an ETDRS level should belong in the moderate or severe category, with the final placement in moderate disease. I find this more informative than a clean table of final labels. It shows that category boundaries result from a judgment about evidence and intended use. A disagreement near a boundary may consequently reflect the difficulty of the operational definition, rather than a model or reader failing to perceive an obvious biological separation.

What the paper licenses is use of a shared terminology with a documented rationale. It does not, by itself, establish how reliably every future reader will apply that terminology, how accurately a particular photograph depicts the required findings, or how well a model will transfer between grading programs. Agreement among experts about a definition and agreement among graders applying that definition are different measurements. The latter needs its own data and reading protocol.

The wording “no apparent retinopathy” is particularly useful to retain. An observation made under specified examination conditions has a narrower meaning than an unrestricted claim of absence. In an image dataset, a negative label may reflect a sufficiently informative negative examination, limited image coverage, or an unsuccessful attempt to see the relevant findings. These situations should not be silently merged. Otherwise, a model may learn the acquisition or documentation process that produced the label.

The ordinal structure also limits what a performance metric means. The grades are ordered, but their numerical coding does not establish equal clinical distances between adjacent categories. An average absolute grade error treats a one-step difference consistently as a numerical event, while the consequence of that difference can depend on where it occurs. I would therefore inspect the full confusion matrix and errors around the decision boundary relevant to the proposed use. A single agreement statistic can conceal the disagreement that matters most.

For a retinal AI dataset, the useful documentation would specify the observation unit, image fields, permitted clinical information, reader qualifications, adjudication procedure, and handling of ungradable images. It should also say whether the target describes an individual image, an eye, or a patient-level decision. Assigning an examination-level grade to every saved image can create an evidence mismatch: the diagnosis may be well established even when a particular frame does not show the finding that justifies it.

Ungradable cases deserve their own evaluation denominator. Excluding them may be reasonable for a narrowly defined image-classification experiment, but it changes the population to which the reported result applies. A screening workflow must also decide what happens when adequate grading is impossible. The model’s ability to recognize insufficient evidence and route the case appropriately is distinct from its ability to classify images already judged gradable.

Reader disagreement should be preserved long enough to understand its source. An adjudicated grade provides a convenient final target, but it can hide whether the initial disagreement concerned lesion identification, severity thresholds, or image quality. Soft labels or multiple-reader targets may represent some of that uncertainty, although they also change the learning objective. The important point is to make that choice explicit rather than treating a consensus label as if it had no production history.

This paper supports Dataset Design, Ground Truth, and Reference Standards because it makes the reference process visible. It also gives Label Quality and Interobserver Variability a concrete distinction to preserve: consensus on a taxonomy, reproducibility in applying it, and validity for the intended clinical claim require different evidence. Improving one does not settle the others.

The connection to Evaluation Beyond AUROC is the need to define the action before collapsing grades into a binary endpoint. A threshold chosen for one workflow may not represent another. In my gallbladder work, the transferable lesson is to document the clinical construct before treating labels as targets. I would want every severity or morphology label to carry its observation conditions, uncertainty, and intended interpretation, so that model agreement remains a claim I can actually defend.