ICSPIS 2025 Oral

KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification

Erfan Nourbakhsh1, Nasrin Sanjari2, Ali Nourbakhsh3

¹University of Isfahan
²Shahid Beheshti University of Medical Sciences
³Isfahan University of Technology

11th International Conference on Signal Processing and Intelligent Systems (ICSPIS 2025)
Oral Paper

Paper

Abstract

Age-related macular degeneration (AMD) and choroidal neovascularization (CNV)-related conditions are leading causes of vision loss worldwide, with optical coherence tomography (OCT) serving as a cornerstone for early detection and management. However, deploying state-of-the-art deep learning models like ConvNeXtV2-Large in clinical settings is hindered by their computational demands. Therefore, it is desirable to develop efficient models that maintain high diagnostic performance while enabling real-time deployment. In this study, a novel knowledge distillation framework, termed KD-OCT, is proposed to compress a high-performance ConvNeXtV2-Large teacher model, enhanced with advanced augmentations, stochastic weight averaging, and focal loss, into a lightweight EfficientNet-B2 student for classifying normal, drusen, and CNV cases. KD-OCT employs real-time distillation with a combined loss balancing soft teacher knowledge transfer and hard ground-truth supervision. Evaluated on the Noor Eye Hospital (NEH) dataset with patient-level cross-validation, KD-OCT achieves near-teacher performance with substantial reductions in model size and inference time, facilitating edge deployment for AMD screening. Code: github.com/erfan-nourbakhsh/KD-OCT.

Background

SOTA Block Architectures

Teacher builds on ConvNeXtV2 with global response normalization. Goal: keep that accuracy in a much smaller student.

Comparison of Swin, ResNet, ConvNeXt, and ConvNeXtV2 block architectures.

Figure 1. Comparison of block architectures in SOTA models for medical image analysis: (a) Swin Transformer Block; (b) ResNet Block; (c) ConvNeXt Block; (d) ConvNeXtV2 Block with global response normalization (GRN).

Method

Data Prep and Augmentation

Patient-level splits avoid leakage. Training uses heavy RandAugment; validation stays minimal; inference uses 5-variant TTA.

Overview of data preparation splits for NEH and UCSD.

Figure 2. Overview of data preparation.

Training, validation, and TTA augmentation pipelines.

Figure 3. Overview of the data augmentation pipelines in KD-OCT, including the training sequence with RandAugment and geometric/color transforms, minimal validation steps, and Test-Time Augmentation (TTA) variants for inference.

Teacher model training overview with focal loss and SWA.

Figure 4. Overview of the teacher model training.

Distillation

KD-OCT Framework

Real-time distillation from ConvNeXtV2-Large to EfficientNet-B2. Soft KL (α=0.7, T=4) plus hard CE (β=0.3).

KD-OCT framework transferring knowledge from ConvNeXtV2-Large teacher to EfficientNet-B2 student.

Figure 5. Overview of the KD-OCT framework, showing knowledge transfer from the ConvNeXtV2-Large teacher to the EfficientNet-B2 student via real-time distillation.

Results

NEH Three-Class Performance

Patient-level 5-fold CV on NEH (normal / drusen / CNV). Student matches teacher accuracy with 25.5× fewer parameters (196.4M → 7.7M).

Model Param (M) Accuracy Sensitivity Specificity
VGG16*28.391.6 ± 2.291.4 ± 2.095.6 ± 1.1
ResNet50*23.686.8 ± 2.086.4 ± 1.693.0 ± 0.9
DenseNet121*7.090.0 ± 1.489.7 ± 1.794.7 ± 0.8
EfficientNetB0*4.085.4 ± 2.684.5 ± 2.292.1 ± 1.3
FPN-VGG16*21.692.0 ± 1.691.8 ± 1.795.8 ± 0.9
FPN-DenseNet121*14.390.9 ± 1.490.5 ± 1.995.2 ± 0.7
SF net29.282.6 ± 2.480.4 ± 2.896.2 ± 0.6
MedSigLIP430.484.5 ± 3.281.81 ± 4.6494.42 ± 1.09
KD-OCT Teacher (ConvNeXtV2-L)196.492.6 ± 2.392.9 ± 2.198.1 ± 0.8
KD-OCT Student (EfficientNet-B2)7.792.46 ± 1.3692.15 ± 1.2996.04 ± 0.78

Table 1. Three-class NEH results with five-fold patient-level CV. *Reported from prior work. Bold: best in KD-OCT rows / compact param count.

Transfer

UCSD Four-Class Generalization

Without fine-tuning, teacher and student both reach 98.4% on the UCSD test set (normal / drusen / CNV / DME).

Model Accuracy Sensitivity Specificity
Kaymak et al.*97.198.499.6
Hassan et al.* (w/ preprocess)98.698.2799.6
FPN-VGG16*98.410097.4
KD-OCT Teacher98.498.4599.47
KD-OCT Student98.498.4099.47

Table 2. UCSD four-class test-set results (no preprocess for KD-OCT).

Conclusion

Main Takeaways

BibTeX

@INPROCEEDINGS{11551784,
  author={Nourbakhsh, Erfan and Sanjari, Nasrin and Nourbakhsh, Ali},
  booktitle={2025 11th International Conference on Signal Processing and Intelligent Systems (ICSPIS)},
  title={KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification},
  year={2025},
  pages={605-611},
  doi={10.1109/ICSPIS68676.2025.11551784}
}