threshold policy design

Designs and analyzes threshold-based decision rules that map the current state (e.g., price and remaining time) to a buy-or-wait action; work includes deriving time-dependent purchase thresholds, formulating and solving the differential/recursive equations that govern threshold evolution (ODEs or dynamic-programming recursions), and computing optimal stopping thresholds for implementation and policy evaluation.

thresholdpolicydesign

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Dynamic Pricing of an Expiring Item under Strategic Buyers with Stochastic Arrival

Sep 29, 2025
SC
Suyeon Choi
🏛️ KAIST | Seoul National University

This paper studies optimal dynamic pricing for time-sensitive goods (e.g., expiring vouchers) in a setting with strategic buyers possessing private valuations and stochastic arrival times. The central challenge lies in the strategic tension between the seller’s urgency to liquidate inventory before expiration and buyers’ incentive to delay purchase in anticipation of price reductions. To address this, we propose a Value-Based Threshold (VBT) policy that decouples the buyer’s two-dimensional private type—valuation and arrival time—and rigorously establish existence and constructive characterization of the Bayesian Nash equilibrium. Leveraging stochastic processes, dynamic game theory, and ordinary differential equations, we develop an analytically tractable equilibrium framework. Numerical analysis reveals: (i) linear discounting is near-optimal in thick markets, whereas fixed pricing dominates in thin markets; (ii) high buyer time-sensitivity favors linear markdowns, while patient sellers benefit from a quasi-auction mechanism—charging premium prices early and steeply slashing prices near expiration. Our results fundamentally characterize how market thickness and seller patience jointly determine optimal pricing structure.

Developing tractable strategies for two-dimensional buyer typesOptimizing dynamic pricing for expiring itemsResolving seller urgency versus buyer waiting incentives

The Support and Resistance Line Method: An Analysis via Optimal Stopping

Mar 03, 2021
VH
Vicky Henderson
🏛️ University of Warwick

Technical analysis lacks a rigorous mathematical foundation for identifying support and resistance levels, leading to ad hoc and subjective trading rules. Method: We propose a path-dependent three-state stock price model and formulate timing decisions for buying and selling as a coupled optimal stopping problem. Employing probabilistic methods, optimal stopping theory, and multi-state Markovian dynamics, we analyze various price processes under heterogeneous risk-aversion preferences. Contribution/Results: We establish, for the first time, the C¹-continuity of the value function with respect to the scale function and develop a unified framework for jointly solving the dual free boundaries corresponding to support and resistance levels. Our approach yields closed-form or semi-analytical optimal trading strategies. The results significantly enhance the mathematical rigor, theoretical interpretability, and computational tractability of technical trading rules.

Determine best buy/sell times via linked free boundary problemsModel stock price transitions between three path-dependent statesSolve optimal stopping problems for general reward functions and dynamics

This paper addresses how dynamic information design influences agents’ stopping times and action choices in optimal stopping problems, particularly when the principal lacks intertemporal commitment power to induce dynamically consistent stopping behavior. Method: We develop a unified framework that jointly models dynamic persuasion and optimal stopping, grounded in game theory, Bayesian updating, dynamic programming, and stopping-time theory. Contribution/Results: We prove that, for any agent preference structure, there exists an optimal information structure that guarantees dynamic consistency without commitment. This structure endogenously determines the optimal stopping time, state-contingent actions, and the path of information revelation. Our framework provides both a rigorous theoretical foundation and a computationally tractable design paradigm for information manipulation in sequential decision-making contexts—including algorithmic recommendation systems and regulatory interventions.

It applies the framework to dynamic persuasion problems with binary/continuous statesIt develops revelation principles for persuasion with and without commitmentThe paper studies how principals influence agent timing and actions through information

This study investigates the performance trade-offs of data-driven approaches in finite-horizon dynamic pricing, with a focus on complex settings involving high-dimensional multi-product offerings, heterogeneous demand structures, and intertemporal revenue constraints. By systematically comparing Fitted Dynamic Programming (Fitted DP) against several reinforcement learning algorithms—integrating demand estimation, trajectory sampling, and expectation-based optimization—the work comprehensively evaluates their relative strengths in terms of revenue generation, stability, constraint satisfaction, and computational scalability. The findings reveal that Fitted DP exhibits superior scalability in structured, complex environments, whereas reinforcement learning demonstrates greater flexibility and adaptability. These insights provide both theoretical grounding and empirical evidence to inform method selection for real-world dynamic pricing systems.

Demand EstimationDynamic PricingDynamic Programming

Identification and Estimation of Continuous-Time Dynamic Discrete Choice Games

Nov 04, 2025
JR
Jason R. Blevins
🏛️ The Ohio State University

This paper addresses identification and estimation in continuous-time dynamic discrete-choice games, focusing on the previously overlooked challenges of endogeneity and heterogeneity in decision arrival rates (i.e., timing of actions). Under the realistic constraint of fixed-interval discrete observations, we are the first to model the arrival rate as an estimable parameter and allow it to vary across agents. Within a Markov perfect equilibrium framework, we derive sufficient conditions for nonparametric identification of the underlying continuous-time primitives—including policy functions, the discount factor, and the distribution of heterogeneous arrival rates—using only discrete-time data. Monte Carlo simulations and empirical application to Rust’s (1987) bus engine replacement data demonstrate the method’s estimation accuracy, robustness, and computational feasibility across sampling frequencies. Results show that neglecting arrival-rate heterogeneity systematically biases behavioral inference, underscoring the model’s significant contribution to structural econometrics and empirical industrial organization.

Establishes equilibrium conditions for generalized modelExamines estimator properties with discrete time dataIdentifies move arrival rates in dynamic choice games

Latest Papers

What's happening recently
View more

This study investigates how decision-makers dynamically choose what to explore and when to stop before taking action. By reformulating the dynamic exploration–stopping problem as a static optimization with an information budget constraint, the authors introduce an information shadow price and establish that the optimal policy exhibits a concavified structure of stopping gains net of this shadow price. They precisely characterize, for the first time, the set of attainable joint distributions and uncover the critical role of time preference curvature in shaping exploration: convex preferences induce Poisson-like exploration, whereas concave preferences restrict stopping to specific time windows. The framework also elucidates the emergence of pure exploration phases and is unifiedly applied to real options, speed–accuracy trade-offs, and continuous-time exploration contests, clearly delineating the structural impact of time preferences and information costs on exploration behavior.

dynamic decision-makingexplorationinformation acquisition

This study addresses the pervasive issue of dynamic inconsistency in classical statistical decision rules based on ex ante criteria—such as minimax regret—which often become irrational to follow once data are observed. The paper provides the first systematic characterization of this problem, develops a formal framework for analyzing dynamic consistency, and axiomatically introduces two novel optimality criteria. By integrating decision theory, axiomatic reasoning, and minimax regret analysis, the proposed criteria effectively prevent post-data deviations across a range of empirical settings, thereby ensuring that decision rules remain coherent before and after information updating. This approach rectifies the behavioral discrepancies inherent in existing methodologies and enhances their practical applicability.

dynamic consistencyeconometricsex ante optimality

This study addresses optimal control in restart-type partially observable Markov decision processes (Restart POMDPs), where the system state is observable only at restart epochs. By introducing a sufficient statistic—comprising the last observed state and the time elapsed since the most recent restart—the problem is transformed into a fully observable MDP. Under both discounted and average cost criteria, the authors establish that when the one-stage cost satisfies a deterioration condition and the transition kernel exhibits monotonicity, the optimal policy admits a time-threshold structure, with thresholds non-increasing in the observed state. The analysis leverages dimensionality reduction via sufficient statistics, stochastic monotonicity, geometric ergodicity, vanishing discount techniques, and uniform boundedness of the relative value function. These results provide a theoretical foundation and a structured policy representation for efficiently solving Restart POMDPs.

Markov decision processoptimal controlPOMDP

Hot Scholars

NA

Nail Akar

Professor of Electrical and Electronics Eng. Dept., Bilkent University
Computer networksperformance evaluationqueuing theorystochastic models
XT

Xiaoqi Tan

Department of Computing Science, University of Alberta
Online algorithmsAlgorithmic economicsDecisions under uncertaintyPerformance evaluation
VP

Vianney Perchet

Crest, ENSAE & Criteo AI Lab
Game TheoryMulti-armed BanditMachine Learning