π€ AI Summary
This work addresses the high token consumption, latency, energy usage, and low accuracy often incurred by generative AI in code editing due to unstructured feedback. To overcome these limitations, the authors propose FileMarkβa structured, line-anchored feedback mechanism that replaces conventional holistic prompting. Through a VS Code extension, multi-model controlled experiments, and function-level automated patch application, the study systematically demonstrates for the first time that line-anchored feedback substantially reduces token usage (by 22%β58%, and up to 24%β80% on large files) while significantly improving correction accuracy, especially for weaker models (gains of +5β7 percentage points, with nearly threefold improvement on large files). The findings further reveal that shifting the editing burden to the feedback mechanism can amplify these benefits.
π Abstract
Generated tokens are a direct driver of the cost, latency, and energy of generative AI (GAI) code editing. We show the format of feedback is a lever on all three. We compare two deliveries of the same requested changes: a holistic prompt (control) versus the structured, line-anchored export of FileMark (treatment). FileMark is a VSCodium extension for inline comments on any file. In a paired experiment line anchoring cut generated tokens by 22% (Claude Opus) and 58% (Claude Sonnet), reaching 24%-80% on files of 100 lines or more, with four of seven models generating significantly fewer tokens after multiple-testing correction. Correctness rose where models had headroom: +2.0 points pooled and +5 to +7 points for three of five local models. An exploratory experiment in which the harness, not the GAI model, applies function-level patches shows the correctness benefit grows further when the edit-application burden is lifted: local-model correctness on 100+ line files roughly triples under anchoring. Line-anchored feedback reduces what stronger models spend and improves what weaker models get right.