Q323 : Few shot image classification on miniImageNet
Thesis > Central Library of Shahrood University > Computer Engineering > MSc > 2026
Authors:
Abstarct: Despite the remarkable success of deep convolutional neural networks in various computer vision tasks, these models inherently depend on massive amounts of labeled training data. In many practical applications, collecting such extensive datasets is challenging, costly, or institutionalized as impossible. Few-shot learning has emerged as a promising paradigm to address this limitation, enabling models to recognize novel visual concepts given only one or a few training examples. However, conventional metric-baxsed methods typically rely on a single feature space extracted from the final laxyer of the network, thereby overlooking the valuable information embedded within shallow and intermediate laxyers. This limitation leads to a severe degradation in performance when encountering complex background noise and domain shifts. In this thesis, a novel and integrated multi-scale metric-baxsed architecture is proposed to enhance the performance of few-shot image classification. The proposed model utilizes a pre-trained ResNet-18 as the feature extraction backbone. Unlike traditional approaches, feature maps are hierarchically extracted from five distinct computational stages to capture a well-balanced blend of fine-grained local structural details and global semantic abstractions. To enrich the feature representations and suppress irrelevant background noise, each of the five feature spaces independently passes through a spatial self-attention module. This module dynamically focuses on the most discriminative regions of the target object to generate refined feature vectors. Consequently, class prototypes are constructed separately across all five feature spaces, and the distances to the query image are measured individually. Finally, a task-adaptive distance fusion module blends these multi-scale distance values using learnable weights, generating the final prediction via a softmax function. To evaluate the efficiency of the proposed model, extensive experiments were conducted on two benchmark datasets, namely miniImageNet and FC100. The empirical results demonstrate that the proposed model achieves highly competitive performance in both 1-shot and 5-shot scenarios, securing an accuracy of 83.15% on miniImageNet and 62.85% on FC100 in the 5-shot classification task, outperforming several state-of-the-art methods. Furthermore, ablation studies validate that the synergy between multi-scale feature extraction and the attention mechanism significantly enhances the model's generalization capabilities and effectively mitigates overfitting when dealing with unseen semantic categories.
Keywords:
#Keywords: Few-Shot Classification #Metric Learning #Multi-Scale Features #Self-Attention #Distance Fusion #Residual Network. Keeping place: Central Library of Shahrood University
Visitor:
Visitor: