Small Language Models for Smart Data Model Classification at the Edge: A Cost-Aware Hybrid Approach

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge of deploying intelligent data classification models on resource-constrained IoT edge devices by systematically evaluating the performance of small language models (SLMs) for this task. The research compares general-purpose, reasoning-oriented, and code-specialized architectures, benchmarking them against zero-cost baselines including TF-IDF and lightweight sentence encoders. Building upon these evaluations, a cost-aware hybrid approach is proposed to optimize the trade-off between classification accuracy and computational efficiency. This work bridges a critical gap in resource-efficient classification solutions for edge computing scenarios, providing essential guidance and practical insights for model selection, task formulation, and deployment strategies on constrained platforms.
πŸ“ Abstract
The rapid proliferation of heterogeneous data sources within the Internet of Things (IoT) across domains such as smart cities, energy management, and environmental monitoring necessitates efficient and scalable data standardization methods. Effective classification of smart data models (SDMs) is essential for facilitating interoperability. However, existing approaches are often limited by high resource consumption and lack applicability in edge environments with constrained computational capabilities. Aiming to bridge this gap, the proposed study evaluates the performance of lightweight open-source language models (LMs) to resolve an input data entity against its corresponding best fitting SDM representation under resource-constrained conditions. It systematically benchmarks a diverse array of models, including general purpose (GP), reasoning-specialized (RS), and code-specialized (CS) architectures, across multiple domain-specific datasets. Addressing the current omission of lightweight, resource-efficient solutions in the literature, the investigation provides significant and valuable insights into model selection, task formulation, and deployment strategies that optimize accuracy and efficiency. A complementary experiment also compares the surveyed large language models (LLMs) against two near-zero-cost similarity baselines (Term Frequency-Inverse Document Frequency (TF-IDF) and a lightweight sentence encoder) on the same task, providing a strong reference point for interpreting the practical value of LLM-based classification on edge platforms.
Problem

Research questions and friction points this paper is trying to address.

Smart Data Model Classification
Edge Computing
Small Language Models
Internet of Things
Resource-Constrained
Innovation

Methods, ideas, or system contributions that make the work stand out.

Small Language Models
Edge Computing
Smart Data Model Classification
Cost-Aware Hybrid Approach
Internet of Things
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
C
Cristian Martella
Dept. of Engineering for Innovation, University of Salento, Via Per Monteroni, Lecce, 73100, Italy
A
Angelo Martella
Dept. of Engineering for Innovation, University of Salento, Via Per Monteroni, Lecce, 73100, Italy
Antonella Longo
Antonella Longo
Associate Professor, University of Salento
Big Data ManagementCyber Physical Social SystemsSmart Cities and citizen scienceService com
Motaz Saad
Motaz Saad
UniversitΓ  del Salento
Big DataNatural Language Processing