Impact Factor 2025:
0.9
CiteScore 2025:
SJR 2025:
Professor of School of Engineering, Design and Built Environment, Western Sydney University, Australia. His research interests cover Industry 4.0, Additive Manufacturing, Advanced Engineering Materials and Structures (Metals and Composites), Multi-scale Modelling of Materials and Structures, Metal Forming and Metal Surface Treatment.
2026-06-18
2026-06-04
2026-08-17
Manuscript received April 7, 2026; revised May 22, 2025; accepted August 19, 2026; published October 8, 2026
Abstract—This paper investigates lightweight Vision Trans- former architectures for few-shot bearing fault diagnosis un- der a leakage-free, held-out-bearing evaluation protocol. Using the Paderborn University bearing dataset at a fixed operating condition (900 rpm, 0.7 Nm, 1000 N), we convert raw 64 kHz vibration signals into filter-bank spectrograms and compare three transformer feature extractors—SimpleViT, ParallelViT, and MobileViT— within a Prototypical Network in a one-shot setting. To avoid the optimistic bias common in small-sample studies, the support and query images are disjoint and the test set consists of bearings never seen during training; accuracy is reported as the mean over 2000 episodes with 95% confidence intervals. Under this protocol, the lightweight MobileViT attains 80.5% ± 0.7% accuracy using only 2.0M parameters, statistically matching SimpleViT (79.9% ± 0.6%, 53.5M parameters) while training about five times faster than ParallelViT (76.7M parameters), which generalizes least well (67.5% ± 0.8%). We further probe the three models under test-time perturbation and find that clean accuracy alone is a poor guide to deployment: all three remain usable under ±10% time jitter, with MobileViT losing 5.7 points and the two pure transformers statistically unchanged, but the MobileViT embedding contracts to a nearly constant vector under additive sensor noise and its accuracy falls to chance, whereas the two pure transformers degrade gracefully. We trace this to the normalization layers rather than to model size or to the hybrid design: MobileViT is the only one of the three that uses BatchNorm, whose running statistics are invalidated by the input shift, and re-estimating those statistics on unlabelled noisy data recovers 26.5 ± 2.4 accuracy points without retraining. Parameter count is therefore not the primary driver of generalization, and normalization choice, not compactness, governs robustness in this setting. Keywords—vision transformers, few-shot learning, bearing fault detection, machine learning, prototypical networks, model efficiency, robustness, normalization layers Cite: Farhan Md. Siraj, Syed Tasnimul Karim Ayon, Mahe Zabin, and Jia Uddin, "A Few-shot Learning Based Light-weight Deep Learning Model for Detecting Bearing Crack Size Using Vibration Sensor Data," International Journal of Mechanical Engineering and Robotics Research, Vol. 15, No. 5, pp. 514-528, 2026. doi: 10.18178/ijmerr.15.5.514-528Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).