BI-RADS-Net: An Explainable Multitask Learning Approach for Cancer Diagnosis in Breast Ultrasound Images

Paper. BI-RADS-Net: An Explainable Multitask Learning Approach for Cancer Diagnosis in Breast Ultrasound Images. Paper manuscript

BI-RADS-Net asks whether breast ultrasound classification can be accompanied by predictions expressed in the vocabulary used to describe the lesion. The question matters because a benign or malignant output leaves a clinician with little to inspect when the model disagrees with an interpretation. Shape, orientation, margin, echo pattern, and posterior features provide more specific claims that can be checked. I read the paper alongside concept-bottleneck work to understand what such claims reveal about the diagnostic computation.

The architecture shares a feature extractor across several tasks. Classification branches predict the five BI-RADS descriptors and tumor class, while a regression branch predicts malignancy likelihood. The clinical motivation is coherent: related supervision might improve the representation and expose recognizable findings. However, that motivation contains two hypotheses. One concerns whether auxiliary tasks improve prediction. The other concerns whether their outputs explain the final decision. The architecture and experiments need to be read separately for each.

The evaluation combined 1,192 tumor-containing images from BUSIS and BUSI, comprising 562 and 630 images respectively. The paper used five-fold cross-validation and an ImageNet-initialized VGG-16 backbone for its main model. This is a combined-dataset development experiment, rather than a demonstration of transfer to an institution excluded from development. The methods describe image splits; they do not establish enough patient-level detail for me to verify independence of all related images across folds.

The most direct comparison adds clinical tasks to a tumor-classification branch. Reported tumor accuracy increased from 86.4% for the single-branch model to 88.9% for the complete model, and all five descriptor tasks exceeded 80% accuracy. These results support the usefulness of the multitask training recipe on the evaluated data. They do not identify how much of the improvement comes from each clinical finding in each prediction, or establish that the same improvement would survive a different acquisition distribution.

The paper also examines preprocessing, pretraining, augmentation, and backbone choices. A detail worth preserving is that some ablations are cumulative. Removing a later component after earlier components have already been removed does not estimate that component’s isolated contribution to the full model. I would read those rows as comparisons between particular pipelines. Treating every difference as an independent causal effect would give the table more explanatory power than its design supports.

The main architectural limitation can be written simply. With shared features \(z\), the descriptor prediction can be \(c=g(z)\), while the malignancy prediction is \(y=h(z)\). The final classifier receives the shared representation directly. It is therefore free to use information that never appears in the displayed descriptors. Accurate descriptor outputs show that the network can recover those features; they do not establish that those outputs mediate the diagnosis.

This distinction remains important even if the auxiliary tasks improve accuracy. Supervision can change the representation during training without creating an inspectable dependency through the displayed outputs at inference. A model could predict a clinically sensible margin and still base part of its malignancy score on an acquisition cue. The descriptor would be useful as an additional measurement, but presenting it as the reason for the score would require further evidence.

A strict concept bottleneck changes that structure by restricting the diagnostic head to concept values. That restriction would answer a stronger computational question, although it would still leave concept measurement and information leakage to assess. BI-RADS-Net is valuable as a baseline precisely because it helps separate the benefit of clinical supervision from the benefit of restricting the prediction pathway. Comparing the two fairly would require attention to backbone capacity, annotation access, optimization, and the information allowed into each head.

Descriptor accuracy also needs more context than a single percentage. A common descriptor category can dominate accuracy while an uncommon but important category remains unreliable. Confusion matrices, class frequencies, and agreement across readers would help establish whether a concept output is clinically usable. In ultrasound, the reference may itself depend on the chosen view. A feature that cannot be assessed from a saved frame should not automatically receive the same interpretation as a confidently absent feature.

The malignancy-likelihood branch raises a related measurement question. Agreement with its training target does not automatically mean that numerical outputs are calibrated probabilities for a new population. Calibration requires comparison with observed outcomes in the population and workflow of interest. Likewise, agreement between a clinician and the displayed descriptors does not establish that the interface improves decisions. The clinician-facing qualitative assessment discussed as future work remains a separate evidentiary step.

For a useful follow-up, I would evaluate the diagnostic output and each descriptor independently, then test their relationship. Clinically validated image pairs could alter a named finding while preserving other relevant evidence as far as possible. The audit would record both the descriptor response and the malignancy response. Merely editing the displayed auxiliary prediction would be uninformative if that value is not an input to the malignancy head. Any intervention must target the actual deployed computation.

This paper supports Clinical Feature Annotation and Multi-Task Learning by showing a concrete use for structured observations beyond the final diagnosis. It also illustrates the note’s warning that observation, interpretation, and action are different targets. A descriptor needs a defensible annotation protocol before it becomes useful supervision, and successful supervision still leaves its role in diagnosis unresolved.

The architectural comparison belongs directly with Clinical Concepts and Concept-Based Interpretability. BI-RADS-Net exposes recoverable clinical information while allowing a diagnostic bypass. Explanation Faithfulness Versus Plausibility then supplies the evaluation boundary: recognizable and accurate findings can make an interface plausible without fully explaining its score. For my gallbladder models, I would retain multitask prediction as a useful measurement and learning strategy, while making any stronger claim about evidence reliance depend on the actual prediction pathway and targeted tests.