🤖 AI Summary
This study addresses the misclassification problem in vision-language models (VLMs) caused by critical information loss during image resizing at preprocessing interfaces. We propose AliasForge, a framework that formally proves the existence of null-space vectors in fixed-point downsamplers and constructs collision and certification criteria based on lattice theory. By leveraging fixed-point resampling, it generates adversarial examples that yield identical interface states yet opposite labels. Experiments analyzing 20 configurations establish cross-architecture certified pairs, demonstrating indistinguishable decision logic across 18 sample sets. This achieves mathematical falsifiability of interface-level errors, precisely delineates model inference boundaries, and reveals the limitations of existing routing strategies in mitigating such vulnerabilities.
📝 Abstract
Vision-language verifiers and routers must distinguish errors repairable by more reasoning from those caused by visual evidence never reaching the language model. This distinction lacks ground truth because annotators see full-resolution images while models receive preprocessed tensors. We introduce AliasForge to create cases where the relevant fact is provably absent from the interface. Fixed-point resampling makes the pre-rounding resize an exact integer linear map that can send nonzero integer perturbations to zero. Hiding a label-flipping perturbation there produces images with opposite step-correctness labels but bit-identical interface states. Every verifier therefore has the same output law on both members, giving pair-balanced accuracy exactly one half and zero gain from language-side repair. We prove that every fixed-point downscaler has such null vectors and bound their smallest size at the most common ratios, which rules out an 8-bit fit whenever the bound exceeds 255. From the resize configuration alone, a lattice criterion supplies realizable collisions and certifies their absence within the specified construction family. It resolves all 20 screened configurations, 17 as constructible and 3 as non-constructible. We construct certified pairs across three architectures and certify four additional processors, with zero decision-logit gap on all 18 scored pairs and none of the 18 controls. The pairs also screen routers that waste computation on re-attention or further reasoning. On natural items, per-item routing headroom exists, but no tested interface-only router improves over stopping. Our fiber ceiling bounds the headroom recoverable from the interface. Code: https://github.com/KurbanIntelligenceLab/aliasforge.