🤖 AI Summary
This study addresses the challenge of quantifying strategic interactions and off-ball movement value among 22 players in soccer, where existing methods often neglect inter-agent dependencies. We model matches as finite-horizon dynamic games grounded in Markov Perfect Equilibrium (MPE). By decomposing decision states for scalable Q-value estimation and integrating reinforcement learning with feature basis decomposition on J-League tracking data, our approach achieves comprehensive player-wide action valuation. Compared to independent RL baselines, the proposed method accurately captures the context-dependency of off-ball runs and defensive positioning, yielding more context-sensitive action rankings. Furthermore, the team-averaged Q-values exhibit a significant positive correlation with expected goals, effectively enhancing both the accuracy and interpretability of tactical evaluation in professional soccer.
📝 Abstract
Valuing player actions in football requires accounting for strategic interactions among 22 players, including off-ball movements and defensive positioning. Existing reinforcement-learning-based methods commonly aggregate decisions at the team level or estimate player values independently, leaving strategic interdependence among players insufficiently represented. This study proposes an action valuation framework inspired by Markov perfect equilibrium (MPE) for all players. Each possession is modeled as a finite-horizon dynamic game, with each player represented as an autonomous agent whose policy depends on the current game state. MPE is used as a motivating solution concept rather than an exact equilibrium. To improve interpretability, we use Expandable Decision-Making States (EDMS) and decompose the Q-value into a successor-feature basis and a linear reward-weight vector. The value basis is estimated by linear TD initialization followed by nonlinear refinement. Using tracking and event data from 95 J1 League matches, we compare the proposed formulation with an independent reinforcement learning baseline. Because the two formulations define TD errors in different target spaces, TD MSE is used only for within-formulation consistency. With EDMS fixed, the independent baseline assigns the highest value to forward movement in 99.21% of evaluated off-ball states, whereas the most frequent direction under the proposed formulation accounts for 17.63%. Team-level average Q-values show a negative association with season-level expected goals for the baseline and a weakly positive association for the proposed formulation. Qualitative analyses illustrate context-dependent valuations of off-ball movements and defensive positioning. Overall, the proposed formulation produces more context-sensitive action rankings, although the comparison does not isolate the MPE-inspired component.