How Do You Know Breast Screening AI Will Perform in the Real World?
Description
This article explores a study that used the Emory Breast Imaging Dataset (EMBED) to evaluate Lunit INSIGHT® DBT across patients of different races, ethnicities, ages, breast densities, and lesion types.
To truly understand performance, AI needs to be evaluated across diverse patients and lesion types. Researchers used the Emory Breast Imaging Dataset (EMBED) to evaluate Lunit INSIGHT® DBT across patients of different races, ethnicities, ages, breast densities, and lesion types. The study included 167,860 screening exams and 1,368 screen-detected cancers. Lunit INSIGHT® DBT achieved an overall AUROC of 0.91, with performance that was not statistically different across race, ethnicity, or age.
Why is evaluating diverse patient populations important?
Real-world screening is not one population.
To truly understand performance, AI needs to be evaluated across diverse patients and lesion types.
What dataset was used in this study?
Researchers used the Emory Breast Imaging Dataset (EMBED) to evaluate Lunit INSIGHT® DBT. The evaluation included patients of different races, ethnicities, ages, breast densities, and lesion types.
The study included:
167,860 screening exams
61,332 women
1,368 screen-detected cancers
How did Lunit INSIGHT® DBT perform?
Lunit INSIGHT® DBT achieved an overall AUROC of 0.91. Performance was not statistically different across race, ethnicity, or age.
What do these findings suggest?
Evidence is strongest when evaluation reflects the diversity of the patients seen every day in screening practice.
This study evaluated Lunit INSIGHT® DBT across a large and diverse patient population and found performance was not statistically different across race, ethnicity, or age.
How many screening exams were included?
The study included 167,860 screening exams.
How many cancers were included in the evaluation?
The study included 1,368 screen-detected cancers.
What was the overall performance of Lunit INSIGHT® DBT?
Lunit INSIGHT® DBT achieved an overall AUROC of 0.91.
Did performance differ across patient groups?
Performance was not statistically different across race, ethnicity, or age.
Why does evaluating diverse populations matter?
The post highlights that evidence is strongest when evaluation reflects the diversity of the patients seen every day in screening practice.