π€ AI Summary
This study investigates whether existing probing methods genuinely capture large language modelsβ awareness of evaluation contexts or merely reflect superficial structural cues from prompt formats. For the first time, it systematically disentangles evaluation context from prompt format by constructing a controlled 2Γ2 dataset and applying diagnostic text rewrites, thereby assessing linear probe effectiveness under partially constrained prompt structures. The results demonstrate that probe signals primarily stem from structural features of benchmark formulations rather than semantic understanding: probe performance substantially degrades when free-form prompts are used. This finding reveals the confounding influence of structural artifacts in current research on model awareness, undermining the reliability of prior conclusions and highlighting the critical need to distinguish genuine semantic comprehension from spurious dependencies on prompt formatting.
π Abstract
Prior work uses linear probes on benchmark prompts as evidence of evaluation awareness in large language models. Because evaluation context is typically entangled with benchmark format and genre, it is unclear whether probe-based signals reflect context or surface structure. We test whether these signals persist under partial control of prompt format using a controlled 2x2 dataset and diagnostic rewrites. We find that probes primarily track benchmark-canonical structure and fail to generalize to free-form prompts independent of linguistic style. Thus, standard probe-based methodologies do not reliably disentangle evaluation context from structural artifacts, limiting the evidential strength of existing results.