🤖 AI Summary
This study addresses whether fine-tuning large language models necessarily depends on human-readable text. To this end, it proposes DASA, a method that discards natural language fluency constraints and directly optimizes data representations for model update utility. By leveraging activation gradient feedback to guide the optimization of continuous synthetic embeddings and integrating LoRA adaptation, DASA replaces conventional text with non-natural-language continuous vectors for efficient fine-tuning. Experimental results demonstrate that the synthetic embeddings generated by DASA achieve fine-tuning performance comparable to or exceeding that of original natural language data. Furthermore, compared to the GRADMM baseline, DASA accelerates training by 3.6× to 4.9× while maintaining similar GPU memory overhead. These findings validate the effectiveness and efficiency of bypassing the discrete text space to directly optimize embedding representations for model adaptation.
📝 Abstract
Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the role of activation gradients in local risk reduction, DASA targets useful adaptation updates rather than source-text reconstruction or linguistic fluency. The resulting embeddings are used directly for downstream fine-tuning; discrete token projections are employed only for qualitative inspection. Experiments on six models from the Llama and Qwen families, ranging from 1B to 32B parameters, cover six benchmarks spanning knowledge, mathematical reasoning, code generation, and commonsense reasoning. Under matched LoRA adaptation settings, DASA achieves performance comparable to the source natural-language data and surpasses it in multiple configurations, while outperforming GRADMM in most comparisons. Further experiments cover general-domain and task-specialized source data. Under the evaluated synthesis settings, DASA provides a $3.6$--$4.9\times$ speedup over GRADMM with comparable peak GPU memory.