🤖 AI Summary
This study addresses the challenge that existing large language model (LLM) fingerprinting techniques struggle to identify the underlying models encoded by intermediaries. To this end, we propose LIDAR, an active black-box fingerprinting approach. Its core innovation lies in extracting identity features from agent execution behaviors rather than textual outputs for the first time, eliminating the need for internal parameter access. Furthermore, LIDAR integrates complementary instance-level and distribution-level feature representations, coupled with a lightweight probabilistic identifier for contrastive analysis. Experimental evaluations across 36 models demonstrate that our method achieves high Top-1 accuracy, significantly outperforming existing baselines.
📝 Abstract
LLMs increasingly operate through coding-agent harnesses that inspect repositories, invoke tools, and modify files. Substituting the model behind such an agent can therefore change security-relevant decisions, including whether it verifies changes or recovers safely from failures. Existing LLM fingerprints largely infer identity from direct text or token distributions. In coding agents, these signals are mediated by system instructions, controller logic, tools, and execution feedback, limiting their transfer.
We present LIDAR (LLM Identification from Decisions and Actions at Runtime), an active black-box fingerprinting method for coding-agent execution. Three coding probe pairs expose post-edit verification, transient-failure recovery, and specification--test conflict resolution under controlled changes. LIDAR represents the resulting trajectories with complementary instance-level and distribution-level features and compares them with clean references using a lightweight probabilistic identifier. It requires no access to model weights, logits, or provider internals.
Across 36 models from seven families and two agent harnesses, LIDAR achieves high Top-1 accuracy and MRR and outperforms four existing fingerprinting and API-auditing baselines. Ablations confirm that the two feature levels, all probe pairs, and their controlled variants contribute. These results show that agent execution behavior provides model-identity evidence beyond final outputs.