🤖 AI Summary
This study addresses the inherent trade-off between privacy leakage and task performance in remote multimodal reasoning by proposing ReCast, an agent plug-in framework. ReCast introduces a novel contract-preserving mechanism that replaces source content via entity rewriting, reversible numerical mapping, and media reconstruction under fixed interfaces, while maintaining task-relevant semantic relationships. To ensure usability, it further incorporates a 4B distilled model for joint rewriting alongside program operand restoration techniques. Experimental results demonstrate that ReCast achieves 75.10% accuracy on benchmarks such as ChartQA, retaining 92.43% of the original performance with a source content leakage rate of only 7.95%. These findings indicate that the proposed framework comprehensively outperforms existing local baselines, offering an effective solution for privacy-preserving multimodal reasoning without compromising downstream task utility.
📝 Abstract
Remote multimodal models offer strong numerical reasoning capabilities over charts and speech, but sending private inputs risks exposing sensitive content. Text-only sanitization cannot directly satisfy fixed media interfaces, while identity anonymization leaves the underlying task content exposed. We introduce ReCast, an agentic plug-in framework that replaces source-specific content while preserving task-relevant relations and the required input modality. ReCast locally converts inputs into a shared textual evidence-query record, jointly rewrites entities and topics with a distilled 4B model, and substitutes values through a locally invertible, role-aware numerical map. A reconstruction agent generates and validates the required media from the protected record. The remote solver returns a program whose protected operands are restored locally before execution. On 4,000 held-out ChartQA and NMSQA examples, ReCast achieves 75.10% accuracy, retaining 92.43% of unprotected remote accuracy, while a model-based audit flags source-content leakage in 7.95% of solver-bound requests. It outperforms all evaluated local baselines, preserving the benefit of remote reasoning while reducing source-content exposure under existing media interfaces.