Adaptive patch sampling and location-aware reasoning for whole body PET-CT multi-organ segmentation
Scientific Reports 16, 22461교신저자 논문
한글 요약
전신 PET-CT 영상에서 여러 장기를 동시에 분할하는 딥러닝 모델의 학습 효율을 높이는 방법을 제안한 연구입니다. 영상을 작은 조각(패치)으로 나누어 학습할 때 기존에는 무작위 또는 고정된 방식으로 조각을 뽑았지만, 이 연구는 모델이 현재 불확실해하거나 오류가 큰 부위에서 더 많이 학습하도록 표본 추출 분포를 스스로 조정하는 적응형 패치 샘플링(APS) 알고리즘을 개발했습니다. 또한 좌표를 직접 주지 않아도 조각이 신체의 어느 위치에서 왔는지 암묵적으로 추론해 특징 표현을 조절하는 패치 인코딩 블록을 함께 도입했습니다. 전신 다장기 PET-CT 분할 실험에서 학습 수렴이 빨라지고 성능이 일관되게 향상되었으며, 외부 데이터셋에서도 견고함이 확인되었습니다. 대규모 의료영상 분할 학습에서 계산 자원을 어디에 쓸지 동적으로 배분하는 전략이 유효함을 보여준 성과입니다.
초록 (English)
Patch-wise learning is a common strategy for training neural networks on large-scale dense prediction problems, yet existing approaches assume uniform or fixed sampling distributions. This assumption is suboptimal when learning difficulty varies spatially and evolves with the model state during optimization. We reformulate patch-wise learning as a dynamic computation allocation problem and propose an adaptive patch sampling (APS) algorithm that learns where to sample by constructing model-state-dependent sampling distributions from voxel-wise uncertainty and prediction error. To learn what contextual information is encoded within sampled patches, we introduce a patch encoding (PE) block that infers implicit location information and modulates feature representations through context-dependent channel-wise attention, without relying on explicit spatial coordinates. Experiments on whole-body multi-organ PET-CT segmentation demonstrate faster convergence and consistent performance gains, with external validation on Synapse dataset confirming robustness. Mechanistic analyses of learning dynamics further characterize sampling behavior induced by APS and representation modulation driven by the PE block through attention analysis and causal channel pruning. Overall, this work contributes an efficient learning strategy for patch-wise training and provides insight into how dynamic sampling and contextual conditioning influence optimization in large-scale dense prediction tasks.