๐ค AI Summary
This work proposes a novel method to quantify changes in key behavioral traitsโsuch as the propensity to acquire sensitive dataโof AI agents during skill evolution. Behavioral traits are modeled as learnable directions in a text embedding space, and a linear model is fitted to the differences between pre- and post-modification embeddings of annotated skill updates to derive generalizable trait vectors. The approach enables evaluation of behavioral tendency shifts for arbitrary modifications via projection onto these vectors. To our knowledge, this is the first framework to support cross-agent assessment of behavioral evolution, facilitating trustworthy intermediaries in scoring skill updates based on trait alignment. Evaluated on 68 annotated instances, the method achieves 91.2% accuracy in sign classification of sensitive data acquisition propensity and a Spearman rank correlation coefficient of ฯ = 0.82.
๐ Abstract
Text files such as skill files, memory files, and behavioral configuration files play a central role in defining how modern agents act. Through edits by humans or the agents themselves, these files may evolve over time, directly steering the agent's behavior in future interactions. We present a methodology and framework for measuring agent $traits$ by defining traits as directions in the embedding space of a text embedding model. We train a linear model on labeled "before" versus "after" skill file diffs to learn a trait vector, then score arbitrary skill edits by projecting their embedding diffs onto this vector. Evaluated on 68 labeled skill diff pairs for the trait of propensity to seek sensitive data, our method achieves 91.2% sign classification accuracy and a Spearman rank correlation of $ฯ= 0.82$ under leave-one-out cross-validation. We build this trait evaluation into a broader agent-to-agent protocol that enables one agent to evaluate another's skill file updates through a trusted intermediary.