Multi-scale spatial attention mechanism-guided feature distillation method for infrared small target detection networks
CSTR:
Author:
Affiliation:

1Research Center for Space Optical Engineering, Harbin Institute of Technology, Harbin 150006;2Institute of Remote Sensing Satellite, China Academy of Space Technology, Beijing 100094;3College of electronic science and technology, National University of Defense Technology, Changsha 410073

Clc Number:

TP753

Fund Project:

Supported by the National Natural Science Foundation of China (61972435, 61401474, 61921001, 62001478)

  • Article
  • |
  • Figures
  • |
  • Metrics
  • |
  • Reference
  • |
  • Related
  • |
  • Cited by
  • |
  • Materials
  • |
  • Comments
    Abstract:

    Existing deep learning methods have achieved significant results in infrared small target detection, but their high computational cost makes them unsuitable for resource-constrained scenarios. There is an urgent need to explore knowledge distillation methods that can balance light weightiness and high accuracy to improve the operational efficiency of infrared small target detection networks. However, due to the extreme characteristics of infrared small targets, conventional distillation methods for infrared detection networks suffer from the loss and diffusion of knowledge about small targets during knowledge transfer and the mismatch of hierarchical feature representations between teacher and student networks. This impairs the student network's ability to learn features of small targets, hindering further improvement in detection capabilities. To address these issues, this paper proposes a feature distillation method guided by a multi-scale spatial attention mechanism. First, a multi-scale spatial attention (MSA) mechanism is designed to capture and fuse multi-scale information of target features, thereby effectively acquiring the target region. Then, an L2 normalization strategy for features is designed to address the differences in feature distribution between teacher and student networks. Finally, an adaptive weighted mean square error (AWMSE) loss function is proposed to guide the student network to strengthen its learning of key target regions. Experimental results on two recognized datasets (NUDT-SIRST, NUAA-SIRST) demonstrate that the proposed distillation method achieves superior detection performance, with the student network even matching the detection performance of the teacher network. Furthermore, the lightweight model after distillation achieves more than 2x inference acceleration when deployed on HUAWEI and NVIDIA edge devices.

    Reference
    Related
    Cited by
Get Citation
Related Videos

Share
Article Metrics
  • Abstract:
  • PDF:
  • HTML:
  • Cited by:
History
  • Received:December 03,2025
  • Revised:February 13,2026
  • Adopted:February 15,2026
  • Online: April 28,2026
  • Published:
Article QR Code