Paper
Age-related macular degeneration (AMD) and choroidal neovascularization (CNV)-related conditions are leading causes of vision loss worldwide, with optical coherence tomography (OCT) serving as a cornerstone for early detection and management. However, deploying state-of-the-art deep learning models like ConvNeXtV2-Large in clinical settings is hindered by their computational demands. Therefore, it is desirable to develop efficient models that maintain high diagnostic performance while enabling real-time deployment. In this study, a novel knowledge distillation framework, termed KD-OCT, is proposed to compress a high-performance ConvNeXtV2-Large teacher model, enhanced with advanced augmentations, stochastic weight averaging, and focal loss, into a lightweight EfficientNet-B2 student for classifying normal, drusen, and CNV cases. KD-OCT employs real-time distillation with a combined loss balancing soft teacher knowledge transfer and hard ground-truth supervision. Evaluated on the Noor Eye Hospital (NEH) dataset with patient-level cross-validation, KD-OCT achieves near-teacher performance with substantial reductions in model size and inference time, facilitating edge deployment for AMD screening. Code: github.com/erfan-nourbakhsh/KD-OCT.
Background
Teacher builds on ConvNeXtV2 with global response normalization. Goal: keep that accuracy in a much smaller student.
Figure 1. Comparison of block architectures in SOTA models for medical image analysis: (a) Swin Transformer Block; (b) ResNet Block; (c) ConvNeXt Block; (d) ConvNeXtV2 Block with global response normalization (GRN).
Method
Patient-level splits avoid leakage. Training uses heavy RandAugment; validation stays minimal; inference uses 5-variant TTA.
Figure 2. Overview of data preparation.
Figure 3. Overview of the data augmentation pipelines in KD-OCT, including the training sequence with RandAugment and geometric/color transforms, minimal validation steps, and Test-Time Augmentation (TTA) variants for inference.
Figure 4. Overview of the teacher model training.
Distillation
Real-time distillation from ConvNeXtV2-Large to EfficientNet-B2. Soft KL (α=0.7, T=4) plus hard CE (β=0.3).
Figure 5. Overview of the KD-OCT framework, showing knowledge transfer from the ConvNeXtV2-Large teacher to the EfficientNet-B2 student via real-time distillation.
Results
Patient-level 5-fold CV on NEH (normal / drusen / CNV). Student matches teacher accuracy with 25.5× fewer parameters (196.4M → 7.7M).
| Model | Param (M) | Accuracy | Sensitivity | Specificity |
|---|---|---|---|---|
| VGG16* | 28.3 | 91.6 ± 2.2 | 91.4 ± 2.0 | 95.6 ± 1.1 |
| ResNet50* | 23.6 | 86.8 ± 2.0 | 86.4 ± 1.6 | 93.0 ± 0.9 |
| DenseNet121* | 7.0 | 90.0 ± 1.4 | 89.7 ± 1.7 | 94.7 ± 0.8 |
| EfficientNetB0* | 4.0 | 85.4 ± 2.6 | 84.5 ± 2.2 | 92.1 ± 1.3 |
| FPN-VGG16* | 21.6 | 92.0 ± 1.6 | 91.8 ± 1.7 | 95.8 ± 0.9 |
| FPN-DenseNet121* | 14.3 | 90.9 ± 1.4 | 90.5 ± 1.9 | 95.2 ± 0.7 |
| SF net | 29.2 | 82.6 ± 2.4 | 80.4 ± 2.8 | 96.2 ± 0.6 |
| MedSigLIP | 430.4 | 84.5 ± 3.2 | 81.81 ± 4.64 | 94.42 ± 1.09 |
| KD-OCT Teacher (ConvNeXtV2-L) | 196.4 | 92.6 ± 2.3 | 92.9 ± 2.1 | 98.1 ± 0.8 |
| KD-OCT Student (EfficientNet-B2) | 7.7 | 92.46 ± 1.36 | 92.15 ± 1.29 | 96.04 ± 0.78 |
Table 1. Three-class NEH results with five-fold patient-level CV. *Reported from prior work. Bold: best in KD-OCT rows / compact param count.
Transfer
Without fine-tuning, teacher and student both reach 98.4% on the UCSD test set (normal / drusen / CNV / DME).
| Model | Accuracy | Sensitivity | Specificity |
|---|---|---|---|
| Kaymak et al.* | 97.1 | 98.4 | 99.6 |
| Hassan et al.* (w/ preprocess) | 98.6 | 98.27 | 99.6 |
| FPN-VGG16* | 98.4 | 100 | 97.4 |
| KD-OCT Teacher | 98.4 | 98.45 | 99.47 |
| KD-OCT Student | 98.4 | 98.40 | 99.47 |
Table 2. UCSD four-class test-set results (no preprocess for KD-OCT).
Conclusion
@INPROCEEDINGS{11551784,
author={Nourbakhsh, Erfan and Sanjari, Nasrin and Nourbakhsh, Ali},
booktitle={2025 11th International Conference on Signal Processing and Intelligent Systems (ICSPIS)},
title={KD-OCT: Efficient Knowledge Distillation for Clinical-Grade Retinal OCT Classification},
year={2025},
pages={605-611},
doi={10.1109/ICSPIS68676.2025.11551784}
}