Institution profile

Tübingen AI Center

Academic institutioneurope · de
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

Jun 24, 2026

This work investigates how to reverse-engineer executable decision code of an agent solely from its behavioral trajectories in games and examines how actively designed adversarial experiments can enhance reconstruction fidelity. To this end, we introduce RevengeBench, a benchmark comprising 75 Elo-calibrated CodeClash strategies across five environments, where customized behavioral probes are generated by observing interactions between target and opponent agents to reconstruct underlying code logic. We formalize strategy reversal as a tractable inverse problem in code space and incorporate a mechanism for controlled experimentation. Experiments across 12 large language models demonstrate that our approach significantly reduces initial behavioral divergence (by 34%–72%) and yields reconstructed strategies that exhibit competitive performance in downstream adversarial settings, particularly bolstering the counterplay capabilities of weaker models.

0 citationsRead paper
Recent publications

Latest Papers

RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments

Jun 24, 2026

This work investigates how to reverse-engineer executable decision code of an agent solely from its behavioral trajectories in games and examines how actively designed adversarial experiments can enhance reconstruction fidelity. To this end, we introduce RevengeBench, a benchmark comprising 75 Elo-calibrated CodeClash strategies across five environments, where customized behavioral probes are generated by observing interactions between target and opponent agents to reconstruct underlying code logic. We formalize strategy reversal as a tractable inverse problem in code space and incorporate a mechanism for controlled experimentation. Experiments across 12 large language models demonstrate that our approach significantly reduces initial behavioral divergence (by 34%–72%) and yields reconstructed strategies that exhibit competitive performance in downstream adversarial settings, particularly bolstering the counterplay capabilities of weaker models.

0 citationsRead paper