belief-dependent experimentation

Design, build, or analyze models, algorithms, and policies for experimentation and selection where agents’ subjective beliefs explicitly shape exploration, choice, and effort—e.g., exploration policies, selection rules, and search procedures that condition on perceived outcomes or correlations. This work includes characterizing equilibrium experimentation strategies, comparing optimistic versus pessimistic belief-driven exploration, and quantifying how belief-dependent selection biases affect individual effort and learning.

belief-dependentexperimentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Learning with Episodic Hypothesis Testing in General Games: A Framework for Equilibrium Selection

Jul 30, 2025
RY
Ruifan Yang
🏛️ Cornell University | UC Berkeley

This paper addresses equilibrium selection in finite normal-form games by proposing a multi-agent learning dynamic grounded in statistical hypothesis testing, designed to converge to a Nash equilibrium that maximizes the minimum (transformed) utility across all players. Methodologically, agents periodically test hypotheses about opponents’ strategies, updating their own policies via empirical observations and belief resampling, while incorporating a utility-dependent exploration decay mechanism to jointly optimize belief refinement and exploration. The key contribution is the first integration of statistical hypothesis testing into game-theoretic learning—enabling endogenous selection of highly robust equilibria without external refinement criteria. Theoretically and empirically, the algorithm converges to an approximate Nash equilibrium set in general finite games and consistently favors solutions that improve the global minimum utility. This provides a novel, interpretable, and adaptive paradigm for equilibrium selection.

Converges to approximate Nash equilibria in general gamesDevelops hypothesis testing-based learning for equilibrium selectionSelects equilibria maximizing minimum utility across players

Exploration Is Not What It Seeks: Catalytic Exploration under Status Quo Uncertainty

Nov 22, 2025
ZH
Zeyu He
🏛️ Tsinghua University | University of California, Berkeley

This paper examines “catalytic exploration”—rational search for alternatives expected to be rejected, undertaken solely to resolve uncertainty about the status quo. Such exploration distorts equilibria in signaling games, reverses standard preferences for information precision (favoring precision about the status quo over alternatives), and generates negative externalities. Method: We develop an options-based decomposition framework distinguishing switch value from catalytic value, integrating Bayesian signal models with game-theoretic analysis. Contributions/Results: (1) We prove that strong catalytic incentives cause collapse of separating equilibria; (2) we identify a counterintuitive “status-quo precision bias” in information acquisition, challenging rational inattention theory; (3) we show that advances in information technology may reduce social welfare by inducing excessive benchmarking. Our results demonstrate that high exploration rates can coexist with low actual switching, underscoring the profound implications of exploratory motives for information architecture and institutional design.

Agents explore alternatives they expect to rejectCatalytic exploration creates negative welfare externalitiesHigh exploration rates coexist with bounded switching

A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments

Sep 29, 2025
MC
Manuel Cherep
🏛️ MIT | Tsinghua University | Dartmouth College

While LLM-powered software agents are increasingly deployed in real-world decision-making domains (e.g., consumer choice, healthcare), existing evaluations predominantly assess task performance, neglecting whether agents exhibit human-like cognitive biases in their decisions. Method: We introduce ABxLab—the first open, behaviorally grounded evaluation framework for AI agent decision-making—employing controlled experiments in a simulated e-commerce environment to systematically manipulate variables such as price, user ratings, and psychological nudges, and quantitatively measure choice biases. Contribution/Results: Our experiments reveal that even without cognitive constraints, agents exhibit statistically significant, human-resembling preference biases—yet their underlying generative mechanisms differ fundamentally from human cognition. ABxLab establishes a novel paradigm for behavioral AI science and serves as a scalable, reproducible benchmark for evaluating agent decision quality beyond functional correctness.

Assessing agent susceptibility to pricing and psychological influencesDeveloping framework to evaluate behavioral biases in AI agentsStudying AI agent decision-making in realistic consumer environments

This work addresses reinforcement learning from sparse, pairwise trajectory preferences over human demonstrations, aiming to efficiently infer an underlying reward function. We propose the first meta-algorithm that simultaneously provides theoretical guarantees and computational tractability: it generates trajectory pairs via randomized exploration, employs batched querying combined with A-optimal experimental design to enable parallel preference elicitation and substantially reduce query complexity, and interfaces with standard RL oracles. In general MDPs, the algorithm achieves an $O(sqrt{T})$ regret bound while ensuring convergence of the final policy. Empirical results demonstrate that the method attains policy performance comparable to explicit reward–driven RL using only a minimal number of preference queries—far fewer than required by conventional approaches—thereby validating both its theoretical rigor and practical efficiency.

Designing algorithms for informative preference queriesImproving query complexity with batch trajectory pairsLearning from human feedback in Markov decision processes

This work addresses the theoretical underpinnings of the exploration–exploitation trade-off in active inference, where balancing curiosity and utility is essential for consistent learning and regret-minimized decision-making. By minimizing expected free energy (EFE), the authors unify information gain and task performance within a single framework and provide the first theoretical guarantees for EFE-minimizing agents: sufficient curiosity simultaneously ensures Bayesian posterior consistency and bounded cumulative regret. The study establishes a formal relationship between the intensity of curiosity and both learning consistency and optimization efficiency, thereby bridging active inference with Bayesian experimental design and Bayesian optimization. Empirical results validate the proposed tuning criterion for curiosity-driven exploration.

Active InferenceBayesian ConsistencyCuriosity

Latest Papers

What's happening recently
View more

This study addresses sequential decision-making in humans concerning persistence versus abandonment, and investigates phenomena such as pessimism traps and insufficient ambition arising from behavioral biases and social influence in social cognition. By constructing an enhanced multi-armed bandit model, the work introduces data-driven algorithmic design into sequential decision-making for the first time and provides a formal characterization of pessimism traps. Theoretical analysis yields nearly tight upper and lower bounds on sample complexity under general conditions, demonstrating that polynomially many samples suffice to learn near-optimal policies. Furthermore, the paper proposes a sustainable community intervention mechanism that effectively disrupts pessimism traps, thereby bridging abstract theories in social epistemology with the complexities of real-world decision-making.

behavioral influencesmulti-armed banditspessimism traps

This work addresses the limitations of traditional Bayesian incentive compatibility, which fails when agents possess private information, lack a common prior, or operate in incomplete information environments. The paper proposes a more general incentive exploration framework that dispenses with Bayesian assumptions and requirements of complete information, allowing agents to act according to any undominated strategy. By introducing a novel definition of incentive compatibility that does not rely on a common prior and integrating multi-prior robust optimization with non-Bayesian decision theory, the framework effectively handles action ties. This approach substantially extends the applicability and robustness of incentive-compatible mechanisms under information asymmetry and uncertainty.

BayesianismCommon PriorFull-Information

This study addresses how intelligent agents can achieve rational decision-making in unknown, non-stationary, and even adversarial environments, and investigates the emergence and stability of equilibria in multi-agent interactions. The work proposes a unified analytical framework that integrates regularized learning strategies, adversarial multi-armed bandit models, and fictitious play dynamics, accommodating both oracle information and bandit feedback settings. Key contributions include deriving optimal regret bounds for single-agent learning in adversarial environments, establishing ergodic convergence to Nash equilibria in zero-sum games, and formulating a folk-theorem-like correspondence between attractors of regularized learning dynamics and Nash equilibria, thereby revealing a fundamental alignment between strategic stability and adaptive learning behavior.

adversarial banditsequilibriumlearning in games

This study addresses failures in conditional reasoning, neglect of relevance, and violations of iterated expectations that arise when agents operate under partial observability and limited inference depth. To account for these cognitive biases, the authors propose a finite belief propagation framework, which formalizes such limitations as asymmetries in reasoning depth for the first time. The model, grounded in directed acyclic graphs, differentiates between inference processes over observed and unobserved variables and integrates both Bayesian and non-Bayesian updating criteria. Validated through behavioral experiments and strategic settings—including public goods provision and social learning—the framework not only explains established findings in conditional reasoning but also demonstrates how bounded inference systematically shapes economic decision-making.

belief updatingcontingent thinkingcorrelation neglect

Hot Scholars

YD

Yifan Dai

Hunan University
LLMAgentAI4Science
YC

Yejin Choi

Stanford University / NVIDIA
Natural Language ProcessingDeep LearningArtificial IntelligenceCommonsense Reasoning
KE

Kee-Eung Kim

KAIST
Machine LearningReinforcement LearningDialogue Systems