learn environment abstractions

Designs and analyzes environment-abstraction mappings and state-aggregation schemes that reduce the complexity of decision problems while aiming to preserve policy quality; builds algorithms to construct or evaluate partitions that enforce shared action distributions within aggregates, separate value-approximation error from action-sharing loss, and provide performance guarantees (bounds) for fixed partitions.

learnenvironmentabstractions

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing environment abstraction methods for large-scale Markov decision processes, which typically prioritize geometric or topological fidelity while overlooking their direct impact on policy performance. The authors propose a policy-performance-oriented tree-structured abstraction mechanism that constructs controllable approximations through state clustering and intra-cluster action distribution sharing. The abstraction granularity is dynamically adjusted based on Q-value discrepancies. Furthermore, the approach explicitly decouples value function approximation error from action-sharing loss and integrates multi-timescale reinforcement learning to enable adaptive refinement and coarsening of the abstract structure. This method achieves substantial state-space compression and demonstrates superior sample efficiency and replanning speed compared to standard actor-critic baselines.

decision-makingenvironment abstractionMarkov decision processes

Homomorphic Mappings for Value-Preserving State Aggregation in Markov Decision Processes

Oct 10, 2025
SZ
Shuo Zhao
๐Ÿ›๏ธ Zhejiang University of Technology | Qingdao University

This paper addresses the policy performance degradation caused by state aggregation in Markov decision processes (MDPs). To preserve optimality under abstraction, we propose a homomorphism-based optimality-preserving abstraction framework. Its core contribution is establishing sufficient conditions for *optimal policy equivalence*, guaranteeing that policies optimized in the low-dimensional abstract MDP remain optimalโ€”or near-optimal within a controllable error boundโ€”in the original MDP. We design Homomorphic Policy Gradient (HPG) and its enhanced variant EBHPG, providing theoretical guarantees on convergence, approximation error bounds, and policy performance lower bounds. The method integrates homomorphic abstraction theory, linear value function approximation, and policy gradient optimization, supported by rigorous error analysis to ensure generalization. Empirical evaluation across multi-task domains demonstrates that our approach achieves a superior trade-off between computational efficiency and policy performance, outperforming seven baseline algorithms.

Balancing computational efficiency with performance loss in aggregationEnsuring optimal policy equivalence between abstract and ground MDPsReducing MDP computational complexity while preserving performance

The CAP theorem imposes a fundamental trade-off among consistency, availability, and partition tolerance in distributed systems, rendering simultaneous strong guarantees impossible. Method: This paper proposes a novel formal framework integrating automata theory with economic incentive mechanisms. It introduces game-theoretic reasoning and economic regulation into state-machine models for the first time, enabling partition-aware modeling and formalizing CAP trade-offs as constrained optimization problems via incentive-augmented global transition semantics. Contribution/Results: The framework transcends classical CAP limitations by guaranteeing both strong consistency and high availability within a bounded error margin ฮต. Experimental evaluation demonstrates that the system maintains convergence, liveness, and correctness under adversarial network partitions. By unifying formal verification with incentive-aligned design, this work establishes a theoretically rigorous and practically deployable foundation for next-generation distributed consensus protocols.

Modeling distributed systems as partition-aware state machinesPreserving availability and consistency within bounded epsilon marginsResolving CAP theorem trade-offs via automata-theoretic economic design

Existing state abstraction methods lack a general principle for rigorously preserving behavioral structure. This work proposes a unified framework that defines behavioral semantics in reinforcement learning in a compositional manner, grounded in local one-step descriptions of system dynamics, and establishes a theory for safe transfer of behavioral structure between abstract and concrete systems. For the first time, the framework enables a compositional formalization of behavioral semantics, supporting the derivation of quantitative metrics with correctness guarantees from logical semantics. It thus lays a principled foundation for behavioral reasoning under state abstraction and provides reusable definitions and provably faithful transfer mechanisms applicable to a broad class of behavioral structures.

behavioral semanticsbehavioral structurescompositional reasoning

Investigating Intra-Abstraction Policies For Non-exact Abstraction Algorithms

Oct 28, 2025
RS
Robin Schmรถcker
๐Ÿ›๏ธ Leibniz University Hannover | University of Southern Denmark

In Monte Carlo Tree Search (MCTS), state/action abstraction often collapses multiple distinct actions into a single abstract node, leading to ambiguous action selection; conventional random tie-breaking yields suboptimal policies and degrades search efficiency. Method: We propose several novel intra-abstraction decision strategies that systematically refine action selection *within* abstract nodes, integrate them into a UCB-based MCTS framework, and synergistically combine them with imprecise abstraction techniques (e.g., pruned Optimistic Graph Abstraction). Contribution/Results: Extensive experiments across diverse benchmark environments and parameter configurations demonstrate that our strategies significantly outperform random tie-breaking baselinesโ€”accelerating convergence of abstraction-guided search, improving policy quality, and enhancing generalization stability. The approach establishes an interpretable, reusable internal decision paradigm for abstraction-augmented MCTS.

Addressing UCB tiebreak issues in pruned OGA algorithmsEvaluating alternative intra-abstraction policies for MCTSImproving MCTS sample efficiency via state abstractions

Latest Papers

What's happening recently
View more

Traditional state abstraction relies on hard partitions, which struggle to effectively model shared interface statesโ€”such as doors or hubsโ€”common in navigation and hierarchical decision-making. This work introduces graph tangles into reinforcement learning for the first time, proposing the tangle-core abstraction framework: it constructs overlapping abstract states via low-order separators of the empirical transition graph and represents shared interfaces using membership kernel functions. The approach enables modeling of overlapping regions, provides value-preserving guarantees, and reveals how hard partitions induce avoidable boundary errors at interfaces. Experiments demonstrate that, across bottlenecked tabular domains, procedurally generated mazes, and MiniGrid environments, the method achieves a superior trade-off between compression and return, while also identifying failure modes when transition topology lacks informative structure.

graph tanglesinterface statesMarkov Decision Processes

This work addresses the high computational complexity and inefficiency of value function approximation in high-dimensional structured Markov decision processes (MDPs). By revealing the low-dimensional geometric structure of decision tessellations induced by optimal policies, the authors propose a boundary-driven policy approximation method that directly learns policy regions rather than value functions. They further introduce a policy loss decomposition mechanism that quantitatively links performance degradation to action boundary errors. Evaluated on inventory control and queue admission tasks, the proposed approach significantly reduces policy error and value gap compared to standard reinforcement learning baselines, achieving faster error convergence and enhanced training stability.

approximate dynamic programmingMarkov decision processesoptimal policy

This work addresses the problem of policy synthesis in Markov decision processes (MDPs) under entropy-based constraints that enforce concentration of state visitation distributions. It formalizes entropy maximization as a policy synthesis objective for the first time, establishes its computational complexity, and introduces a novel method combining convex duality theory with invariant synthesis to handle nonlinear entropy constraints in a conditionally complete manner. By systematically analyzing the roles of memory and randomization in policies, the approach effectively synthesizes and verifies entropy-constrained policies across multiple benchmark instances, substantially extending the expressiveness and applicability of existing policy synthesis frameworks.

concentration propertycontrol policy synthesisentropy objectives

This work addresses the lack of a general mechanism in reinforcement learning for dynamically adjusting the granularity of state-action abstraction, which often leads to a trade-off between task simplification and preservation of critical information. The authors propose an adaptive soft abstraction method grounded in rate-distortion theory, uniquely integrating rate-distortion optimization with bisimulation metrics to enable continuously tunable abstraction over both state and action spaces. By decomposing value function error into learning and abstraction components via performance certificates, they introduce an error-driven dynamic refinement strategy that triggers abstraction optimization when these two error sources become comparable. Empirical results demonstrate that the approach maintains near-optimal performance across multiple tabular environments, even under substantial lossy compression of the state-action space.

dynamic abstractiongranularity adjustmentrate-distortion

This study addresses the problem of extracting an optimal subplan from an existing plan under a budget constraint, while preserving the original actions and their execution order. The goal is to identify a subplan that respects a given cost upper bound, remains executable, and maximizes utility. The decision variant of this problem is proven to be NP-complete. To tackle it, the authors propose a refined integer linear programming (ILP) formulation that significantly reduces model size and enhances computational efficiency without sacrificing solution accuracy. Together with over-subscription planning (OSP), this ILP approach constitutes one of two exact solution methods. Compared to prior work, the proposed ILP method demonstrates marked improvements in both scalability and empirical performance.

action orderingbudget constraintscost-bounded planning

Hot Scholars

MS

Michael S. Bernstein

Professor of Computer Science, Stanford University
Human-computer interactionsocial computinghuman-centered AI
YW

Yangang Wang

Professor, Southeast University
Computer graphicsComputer visionComputational photography
ZM

Ziming Mao

UC Berkeley
Distributed SystemsBig DataAI Systems
XX

Xuan Xie

Macau University of Science and Technology
Trustworthy LLMCyber Physical SystemNeural Network Verification
WD

Wenbin Dai

Shanghai Jiao Tong University
Industrial Edge ComputingIndustrial InformaticsAutomation Code GenerationIndustrial Control Software