Home > Articles > All Issues > 2026 > Volume 15, No. 5, 2026 >
IJMERR 2026 Vol.15(5):514-528
doi: 10.18178/ijmerr.15.5.514-528

A Few-shot Learning Based Light-weight Deep Learning Model for Detecting Bearing Crack Size Using Vibration Sensor Data

Farhan Md. Siraj 1 , Syed Tasnimul Karim Ayon 2, Mahe Zabin 2, and Jia Uddin 3,*
1. Department of Computer Science and Engineering, BRAC University, Dhaka-1212, Bangladesh
2. Human and Digital Interface Department, JW Kim College, Woosong University, Daejeon, Republic of Korea
3. Artificial Intelligence and Big Data Department, Endicott College, Woosong University, Daejeon, Republic of Korea
Email: {farhan.md.siraj, syed.tasnimul.karim.ayon}@g.bracu.ac.bd (F.M.S. & S.T.K.A.);
mahezabin@wsu.ac.kr (M.Z.); jia.uddin@wsu.ac.kr (J.U.)
*Corresponding author

Manuscript received April 7, 2026; revised May 22, 2025; accepted August 19, 2026; published October 8, 2026

Abstract—This paper investigates lightweight Vision Trans- former architectures for few-shot bearing fault diagnosis un- der a leakage-free, held-out-bearing evaluation protocol. Using the Paderborn University bearing dataset at a fixed operating condition (900 rpm, 0.7 Nm, 1000 N), we convert raw 64 kHz vibration signals into filter-bank spectrograms and compare three transformer feature extractors—SimpleViT, ParallelViT, and MobileViT— within a Prototypical Network in a one-shot setting. To avoid the optimistic bias common in small-sample studies, the support and query images are disjoint and the test set consists of bearings never seen during training; accuracy is reported as the mean over 2000 episodes with 95% confidence intervals. Under this protocol, the lightweight MobileViT attains 80.5% ± 0.7% accuracy using only 2.0M parameters, statistically matching SimpleViT (79.9% ± 0.6%, 53.5M parameters) while training about five times faster than ParallelViT (76.7M parameters), which generalizes least well (67.5% ± 0.8%). We further probe the three models under test-time perturbation and find that clean accuracy alone is a poor guide to deployment: all three remain usable under ±10% time jitter, with MobileViT losing 5.7 points and the two pure transformers statistically unchanged, but the MobileViT embedding contracts to a nearly constant vector under additive sensor noise and its accuracy falls to chance, whereas the two pure transformers degrade gracefully. We trace this to the normalization layers rather than to model size or to the hybrid design: MobileViT is the only one of the three that uses BatchNorm, whose running statistics are invalidated by the input shift, and re-estimating those statistics on unlabelled noisy data recovers 26.5 ± 2.4 accuracy points without retraining. Parameter count is therefore not the primary driver of generalization, and normalization choice, not compactness, governs robustness in this setting.

Keywords—vision transformers, few-shot learning, bearing fault detection, machine learning, prototypical networks, model efficiency, robustness, normalization layers

Cite: Farhan Md. Siraj, Syed Tasnimul Karim Ayon, Mahe Zabin, and Jia Uddin, "A Few-shot Learning Based Light-weight Deep Learning Model for Detecting Bearing Crack Size Using Vibration Sensor Data," International Journal of Mechanical Engineering and Robotics Research, Vol. 15, No. 5, pp. 514-528, 2026. doi: 10.18178/ijmerr.15.5.514-528

Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).