Fine-Tuning a 3B-Parameter LLM on a Smartphone: Characterizing Sustained Training

πŸ“… 2026-10-05
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the absence of comprehensive performance characterization and personalized fine-tuning validation for large language models (LLMs) on mobile devices. Leveraging the MLX framework and LoRA techniques, we systematically evaluate the continuous fine-tuning of a 3B-parameter LLM on iPhones, encompassing memory footprint, latency, thermal behavior, and energy consumption analyses, while rectifying an underlying unscheduled kernel defect. This work presents the first full fine-tuning characterization of billion-parameter models on smartphones. Following the kernel fix, training speed improves by 1.47Γ— and energy consumption decreases by one-third, enabling typical user-specific fine-tuning within a single battery charge cycle. Furthermore, the achieved personalization performance is comparable to server-grade baselines.
πŸ“ Abstract
Multi-billion-parameter LLMs now run on phones for inference, and training them on the device would personalize them without user data leaving the phone. Prior work has measured individual training steps of such models on phones, but not complete training runs, and not whether adapters trained on the device improve personalization. We present the first systematic characterization of a multi-billion-parameter LLM fine-tuned on a mobile device, covering memory, per-step time, thermal behavior, and energy. An iPhone 17 Pro can fine-tune a 3B-parameter LLM to a typical user within one battery charge, and the resulting adapters improve personalization as much as adapters trained on a server. Sustained training throttles the phone to about half its initial throughput, and none of the pausing or burst schedules we tested recovers it. Nearly all of each training step is spent in the frozen base model, most of it in the backward pass, which nine of the ten other runtimes we audited do not accelerate. Apple's MLX had a kernel for it that was never dispatched and was incorrect, and our repair, now merged upstream, trains an adapter 1.47x faster on a third less energy. On-device fine-tuning is feasible on current phones, and making it efficient requires runtimes and operating systems to treat training as a first-class workload.
Problem

Research questions and friction points this paper is trying to address.

On-device fine-tuning
Large Language Models
Sustained training
Personalization
Mobile AI
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-device fine-tuning
Large Language Models
MLX kernel optimization
Sustained training characterization
Mobile personalization