What Did a New Automation Bias Study Find About AI in Mammography?
Breast cancer screening, Lunit INSIGHT MMG
Description
A recent paper published in Radiology examined automation bias when using AI in mammography. The study found that reader accuracy improved when AI was correct, but reader sensitivity decreased in cases where AI missed a cancer.
The findings highlight both the benefits of AI and the importance of training, workflow design and threshold selection.
How was the study designed?
In the study, 10 experienced NHS mammography readers reviewed 60 breast screening cases twice: once without AI support and once with AI support.
Researchers also tracked readers' eye movements to understand how they examined the images.
Importantly, these readers did not routinely use AI in practice.
Was this designed to reflect real-world AI performance?
Before looking at the results, it's important to understand that this study was deliberately designed as a stress test.
To make that possible, the researchers used Lunit INSIGHT® MMG to enrich the dataset with a high number of AI errors, including 14 missed cancers and 14 false alarms.
The goal was to explore automation bias, the tendency to rely on the output of an automated system, and understand how radiologists respond when AI gets things wrong.
What happened when AI was correct?
The first findings were encouraging.
When the AI was correct, overall reader accuracy improved, with AUC increasing from 0.57 to 0.72.
What happened when AI raised a false alarm?
Radiologists did not simply follow the AI.
When AI raised a false alarm, readers correctly dismissed those findings, and specificity increased from 21% to 39%, with no downside observed.
What happened when AI missed a cancer?
This was the more challenging finding.
When the AI missed a cancer, reader sensitivity fell from 71% to 39%.
The eye-tracking data also showed readers spent less time visually searching when AI did not flag a finding.
What do these findings mean?
The study suggests radiologists are very good at identifying and dismissing AI false alarms.
The greater risk appears to be the opposite scenario: when AI suggests there is nothing to see, some readers may search less carefully.
This is where threshold selection, workflow design and training become particularly important.
What did the accompanying editorials say?
One of the key themes from the editorials accompanying the paper was that these findings should be viewed as a call for better threshold setting, training and system design.
The editorials framed AI and the reader as a combined system rather than presenting the results as a reason to distrust AI.
FAQ
Did AI improve reader performance?
Yes. When the AI was correct, overall reader accuracy improved, with AUC increasing from 0.57 to 0.72.
Did radiologists follow AI false positives?
No. Readers correctly dismissed AI false alarms and specificity improved from 21% to 39%.
What is automation bias?
In this context, automation bias refers to the tendency to rely on the output of an automated system.
What was the main concern identified in the study?
The greatest concern was not false alarms. It was the reduction in reader sensitivity and visual search time when AI missed a cancer.