Tiny Models, Tough Limits: Benchmarking and Optimizing Small Language Models for Edge Deployment
Book Chapter
Pissinou Makki, A, Lago Enamorado, L, Saleem, C et al. (2026). Tiny Models, Tough Limits: Benchmarking and Optimizing Small Language Models for Edge Deployment
. 683 LNICST 3-22. 10.1007/978-3-032-22500-9_1
Pissinou Makki, A, Lago Enamorado, L, Saleem, C et al. (2026). Tiny Models, Tough Limits: Benchmarking and Optimizing Small Language Models for Edge Deployment
. 683 LNICST 3-22. 10.1007/978-3-032-22500-9_1
Emerging applications—disaster response drones, in-vehicle assistants, and field medical devices—require on-device language intelligence when cloud links are unreliable, privacy is mandatory, and subsecond latency is nonnegotiable. We benchmark seven SLMs (DistilBERT, MobileBERT, ALBERT, MiniLM, Phi-3 Mini, MobileLLaMA and TinyLLaMA) across four mission-aligned use cases (Watchlist Screening, Threat Detection, Document Triage, Multilingual Routing) on five border-relevant datasets (e.g., GTD, FLORES-200). Under controlled edge-like constraints (mobile-class CPU, 1–8 GB shared memory, intermittent networking), we report task quality (accuracy/F1 or ROUGE), batch-1 inference latency, and peak memory, and we introduce a reproducible, edge-budgeted evaluation protocol for security-critical scenarios. We also outline a path to multimodal edge workloads by pairing compact audio/vision encoders with SLM back ends.