Differential Biases and Variabilities of Deep Learning–Based Artificial Intelligence and Human Experts in Clinical Diagnosis: Retrospective Cohort and Survey Study
JMIR Medical Informatics 9(12), e33049교신저자 논문
한글 요약
딥러닝 인공지능과 인간 전문가가 의료영상 진단에서 서로 다른 종류의 편향과 변동성을 보이는지를 귀내시경 영상 진단을 예로 비교한 연구입니다. 세브란스병원 외래에서 수집한 대규모 귀내시경 영상으로 여섯 가지 귀 질환을 분류하는 딥러닝 모델을 만들고, 질환 비율이 균형 잡힌 검사 세트와 실제 유병률처럼 불균형한 검사 세트에서 이비인후과 전문의, 비전문 의사, 여러 딥러닝 모델의 정확도와 일치도를 비교했습니다. 딥러닝 모델은 이비인후과 전문의와 대등하고 비전문 의사보다 훨씬 높은 정확도를 일관되게 보였지만, 흔한 질환에 치우치는 유병률 편향이 있어 데이터 증강 후에도 희귀 질환의 정확도는 낮았습니다. 반면 의사들은 유병률 편향은 적었지만 개인 간 변동이 컸습니다. 두 집단의 강점이 다르므로, 모델이 영상만 보고 흔한 질환에 편향될 수 있음을 염두에 둔다면 이비인후과 전문의가 부족한 상황에서 인공지능이 다양한 숙련도의 의사들을 보조하는 협력적 역할을 할 수 있음을 시사합니다.
초록 (English)
BACKGROUND: Deep learning (DL)-based artificial intelligence may have different diagnostic characteristics than human experts in medical diagnosis. As a data-driven knowledge system, heterogeneous population incidence in the clinical world is considered to cause more bias to DL than clinicians. Conversely, by experiencing limited numbers of cases, human experts may exhibit large interindividual variability. Thus, understanding how the 2 groups classify given data differently is an essential step for the cooperative usage of DL in clinical application.
OBJECTIVE: This study aimed to evaluate and compare the differential effects of clinical experience in otoendoscopic image diagnosis in both computers and physicians exemplified by the class imbalance problem and guide clinicians when utilizing decision support systems.
METHODS: We used digital otoendoscopic images of patients who visited the outpatient clinic in the Department of Otorhinolaryngology at Severance Hospital, Seoul, South Korea, from January 2013 to June 2019, for a total of 22,707 otoendoscopic images. We excluded similar images, and 7500 otoendoscopic images were selected for labeling. We built a DL-based image classification model to classify the given image into 6 disease categories. Two test sets of 300 images were populated: balanced and imbalanced test sets. We included 14 clinicians (otolaryngologists and nonotolaryngology specialists including general practitioners) and 13 DL-based models. We used accuracy (overall and per-class) and kappa statistics to compare the results of individual physicians and the ML models.
RESULTS: Our ML models had consistently high accuracies (balanced test set: mean 77.14%, SD 1.83%; imbalanced test set: mean 82.03%, SD 3.06%), equivalent to those of otolaryngologists (balanced: mean 71.17%, SD 3.37%; imbalanced: mean 72.84%, SD 6.41%) and far better than those of nonotolaryngologists (balanced: mean 45.63%, SD 7.89%; imbalanced: mean 44.08%, SD 15.83%). However, ML models suffered from class imbalance problems (balanced test set: mean 77.14%, SD 1.83%; imbalanced test set: mean 82.03%, SD 3.06%). This was mitigated by data augmentation, particularly for low incidence classes, but rare disease classes still had low per-class accuracies. Human physicians, despite being less affected by prevalence, showed high interphysician variability (ML models: kappa=0.83, SD 0.02; otolaryngologists: kappa=0.60, SD 0.07).
CONCLUSIONS: Even though ML models deliver excellent performance in classifying ear disease, physicians and ML models have their own strengths. ML models have consistent and high accuracy while considering only the given image and show bias toward prevalence, whereas human physicians have varying performance but do not show bias toward prevalence and may also consider extra information that is not images. To deliver the best patient care in the shortage of otolaryngologists, our ML model can serve a cooperative role for clinicians with diverse expertise, as long as it is kept in mind that models consider only images and could be biased toward prevalent diseases even after data augmentation.