Evolution of Few-Shot Image Classification: A Survey from Metric Learning to Large Model Fine-Tuning

Main Article Content

Yuhang Ji

Keywords

FSIC, metric learning, meta-learning, parameter-efficient fine-tuning, cross-domain generalization

Abstract

While deep learning has achieved remarkable success in image recognition, it relies heavily on massive amounts of labeled data. In real-world scenarios like medical imaging, collecting such large datasets is often expensive or impossible. To overcome this critical bottleneck, Few-Shot Image Classification (FSIC) has emerged as a vital domain, enabling models to learn new categories from extremely limited samples. Based on a comprehensive review of recent literature, this paper systematically divides mainstream FSIC methods into three categories: Metric Learning, Optimization/Meta-Learning, and Transfer Learning with Large Model Fine-Tuning. This survey comprehensively evaluates the underlying mechanisms, inherent strengths, and specific weaknesses of each category. Additionally, the paper provides a clear performance comparison to illustrate the evolutionary impact of these different approaches. Finally, the paper deeply discusses remaining technical challenges, specifically detailing issues like cross-domain generalization and computational dependence, and outlines promising future research directions, such as multimodal knowledge fusion, to further advance the field.

Abstract 11 | PDF Downloads 5

References

  • [1] Snell, J., Swersky, K., & Zemel, R. (2017). Prototypical networks for few-shot learning. Advances in neural information processing systems, 30.
  • [2] Finn, C., Abbeel, P., & Levine, S. (2017, July). Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning (pp. 1126-1135). PMLR.
  • [3] Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S.,... & Sutskever, I. (2021, July). Learning transferable visual models from natural language supervision. In International conference on machine learning (pp. 8748-8763). PmLR.
  • [4] Zhou, K., Yang, J., Loy, C. C., & Liu, Z. (2022). Learning to prompt for vision-language models. International journal of computer vision, 130(9), 2337-2348.
  • [5] Khattak, M. U., Rasheed, H., Maaz, M., Khan, S., & Khan, F. S. (2023). Maple: Multi-modal prompt learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 19113-19122).
  • [6] Zhang, R., Fang, R., Zhang, W., Gao, P., Li, K., Dai, J.,... & Li, H. (2021). Tip-adapter: Training-free clip-adapter for better vision-language modeling. arXiv preprint arXiv:2111.03930.
  • [7] Su, M., He, F., Li, G., & Li, F. (2025). PrototypeFormer: Learning to explore prototype relationships for few-shot image classification. Neurocomputing, 640, 130326.
  • [8] Liu, T., Basu, A., Caverlee, J., & Kong, S. (2025). Solving Semi-Supervised Few-Shot Learning from an Auto-Annotation Perspective. arXiv preprint arXiv:2512.10244.
  • [9] Kang, S., Hwang, D., Eo, M., Kim, T., & Rhee, W. (2023). Meta-learning with a geometry-adaptive preconditioner. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 16080-16090).
  • [10] Zhou, L., Shakeri, F., Sadraoui, A., Kaaniche, M., Pesquet, J. C., & Ben Ayed, I. (2025). UNEM: UNrolled generalized EM for transductive few-shot learning. In Proceedings of the Computer Vision and Pattern Recognition Conference (pp. 9665-9675).
  • [11] Yu, T., Lu, Z., Jin, X., Chen, Z., & Wang, X. (2023). Task residual for tuning vision-language models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 10899-10909).