Model Compression using Adaptive Knowledge Distillation
Developed Adaptive Temperature Scaling (ATS), a knowledge-distillation approach that helps smaller neural networks mimic larger models for faster and cheaper deployment on resource-constrained platforms.
ATS achieved up to a 3.63% accuracy gain over the conventional method without additional training cost, with results across CNN architectures including ResNet, VGG, WideResNet, MobileNet, and ShuffleNet on CIFAR-100 and ImageNet data.
