choice modeling

Designs, estimates, and evaluates probabilistic models of individual or aggregate discrete choices and preferences — including discrete, dynamic, and hybrid choice models and parameterized preference distributions — to fit behavioral data, infer attribute importance and latent factors, and predict held-out decisions. Also builds simulators and analyses of model behavior (e.g., systematicity, manipulability, or phase transitions) and generates synthetic preference profiles from parameterized culture models.

choicemodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.13
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Understanding the decision-making process of choice modellers

Nov 03, 2024
GN
Gabriel Nova
🏛️ Delft University of Technology | University of Leeds

This study investigates how methodological choices made by modelers during discrete choice model construction—particularly in noise pollution policy contexts—induce substantial variation in outcomes. Method: We develop the “Serious Choice Modeling Game,” a behavioral experiment platform that systematically tracks modelers’ data exploration, model specification, and interpretation processes, integrating operational logs, descriptive statistics, and multinomial logit modeling for quantitative analysis. Contribution/Results: We find that while data visualization is widespread, missing-data handling is frequently neglected; parsimony preferences and iteration quality significantly affect model fit and simplicity; and substantial heterogeneity exists across modeling strategies applied to identical data. Crucially, willingness-to-pay estimates vary markedly with methodological choices, undermining policy recommendation reliability. This work provides the first systematic, causal evidence linking modeler behavior to outcomes in policy-oriented choice models, offering empirical foundations for enhancing reproducibility and policy robustness of discrete choice analysis.

Analyzing modellers' decisions in handling data and model specificationExploring preference for simpler models despite complex alternativesUnderstanding variability in choice model outcomes due to diverse workflows

This paper addresses the problem of simultaneous discrete (e.g., occupation choice) and continuous (e.g., hours worked) decisions in dynamic models under pervasive unobserved heterogeneity. Methodologically, it proposes the first nonparametric identification framework, combining an EM algorithm with instrumental-variable quantile regression in a two-step estimator. It further integrates Hotz–Miller–style conditional choice probability construction with structural identification theory to circumvent the high-dimensional computational burden inherent in full-solution approaches. The main contributions are threefold: (i) it achieves the first nonparametric identification of dynamic discrete–continuous joint choice models; (ii) it substantially reduces estimation complexity; and (iii) it ensures consistent and robust estimation of both structural parameters and conditional choice probabilities—even in settings with high-dimensional unobserved heterogeneity.

Identifies dynamic models with discrete and continuous choicesIncorporates unobserved heterogeneity in choice modelsReduces computational burden in complex dynamic estimations

Ordered Probabilistic Choice

Apr 01, 2025
CP
Christopher P. Chambers
🏛️ Georgetown University | University of Maryland | Bilkent University

This paper addresses the challenge of identifying heterogeneous individual-level choice behaviors from macro-level aggregate selection data. To this end, it establishes, for the first time, a systematic theoretical linkage between ordered probit choice models and copula theory, mapping individual heterogeneity onto the structural form of copula functions. The authors propose an analytically tractable representation based on extreme-value theory, enabling unique and unbiased identification of both heterogeneity types and their mixing weights. Methodologically, the approach integrates copula modeling, extreme-value function analysis, and structural identification theory to derive a general closed-form extreme-value representation. This framework overcomes key limitations of conventional aggregate modeling—such as loss of behavioral granularity and identifiability constraints—thereby substantially improving the accuracy, interpretability, and structural fidelity of micro-behavioral inference. It introduces a novel paradigm for discrete choice analysis, behavioral econometrics, and multivariate dependence modeling.

Analyze representations using copula theory resultsIdentify micro-level heterogeneity from macro-level dataLink ordered probabilistic choice to copula theory

A Nonparametric Approach with Marginals for Modeling Consumer Choice

Aug 12, 2022
YR
Yanqiu Ruan
🏛️ Singapore University of Technology and Design | National University of Singapore

This paper addresses the challenge of balancing model parsimony and task-specific applicability (e.g., pricing, product assortment optimization) in consumer choice modeling. We propose a nonparametric approach grounded in marginal utility distributions. First, we provide an exact characterization of the choice probability polytope representable by the Marginal Distribution Model (MDM) and its grouped variant (G-MDM). Second, we develop the first nonparametric optimal fitting estimation framework requiring no parametric assumptions. Third, we prove that G-MDM and the Random Utility Model (RUM) are incomparable—neither subsumes the other. Our method enables efficient computation via linear programming validation, mixed-integer convex optimization, and grouped structural modeling. Empirically, it significantly outperforms the multinomial logit model in expressive power, estimation accuracy, and predictive performance, while achieving substantially higher computational efficiency than RUM. Moreover, it provides theoretically guaranteed prediction intervals for choice probabilities over unseen assortments.

Balance tractability and representational power in choice modelingDevelop parsimonious models for consumer choice behaviorEstablish conditions for data consistency with MDM hypothesis

Just Ask Them Twice: Choice Probabilities and Identification of Ex ante returns and Willingness-To-Pay

Mar 06, 2023
RM
Romuald Méango
🏛️ University of Oxford | CESifo | University of Technology Sydney | IZA

This paper addresses the inefficiency and fragility of conventional stated-preference experiments for probabilistic choices, where ex ante expected returns and willingness-to-pay (WTP) estimates rely on multiple choice rounds and strong parametric assumptions—leading to lengthy surveys and low feasibility for ex ante policy evaluation. We propose a nonparametric identification method requiring at most two probabilistic choices per respondent. It imposes no functional-form assumptions on utility and, for the first time, fully identifies both the population distribution of ex ante expected returns and WTP for structured preference objects (e.g., multidimensional job attributes). Theoretical foundations integrate nonparametric identification theory with structured discrete choice modeling. Applied to elite student employment preferences in Côte d’Ivoire, the method robustly identifies a significant upward effect of public-sector jobs on private-sector hiring costs—demonstrating both empirical validity and direct policy relevance.

Enhancing policy evaluation with richer preference identificationEstimating ex ante returns and WTP distributions efficientlyReducing survey length to two choices per individual

Latest Papers

What's happening recently
View more

Traditional parametric discrete choice models struggle to capture complex decision rules under individual heterogeneity, particularly in modeling policy preferences. This study systematically evaluates four machine learning approaches—multinomial logistic regression, generalized additive models, Siamese neural networks, and Gaussian processes—across five behavioral and social science discrete choice tasks. Using both Monte Carlo simulations and real-world energy policy preference data, models are compared via Bayesian Information Criterion (BIC) and predictive accuracy. Results show that semi-parametric and non-parametric models consistently outperform parametric ones; increasing training sample size and choice rule determinism improves performance by 6%–96% and 0%–55%, respectively. On empirical data, the Siamese neural network achieves the best fit (BIC = 13.351). These findings underscore that model selection should be guided by task characteristics and highlight both the promise and limitations of data-driven methods in policy-oriented discrete choice modeling.

choice rulesdiscrete choice modelingindividual heterogeneity

This study addresses the problem of identifying the distribution of latent behavioral types and their choice patterns from aggregate data when only group-level choices are observed and individual types are unobserved. Assuming minimal qualitative prior knowledge, the authors establish the first necessary and sufficient conditions for the identifiability of behavioral types. They characterize cross-type behavioral heterogeneity through an equivalence condition that combines a matching criterion with the algebraic structure of the mapping matrix linking types to aggregate outcomes. Leveraging tools from combinatorics and matrix algebra, the paper elucidates the mechanism by which type-specific behaviors map to aggregate data and demonstrates that, provided the data exhibit sufficient cross-type heterogeneity, both the latent types and their distribution are uniquely identifiable—thereby establishing a theoretical foundation for nonparametric identification in this setting.

aggregate choicebehavioral typeschoice behavior

This study addresses the bias in parameter estimation commonly arising in discrete choice models due to unobserved consideration sets. The authors propose a practical approach that constructs individual-specific consideration sets based on historical choices, enabling consistent estimation within a Logit framework without relying on exhaustive universal set evaluations or subjective self-reports. Theoretically, the paper provides the first rigorous proof that, under homogeneous choice probability assumptions, such consideration sets satisfy sufficient conditions for consistent estimation, while also offering a refined interpretation of the alternative sampling theorem. Methodologically, the approach is validated through Monte Carlo simulations, synthetic Logit-generated data, and large-scale passive behavioral datasets—such as smart card and mobile phone records. Empirical results demonstrate that the proposed method yields consistent and robust parameter estimates under specified conditions, thereby opening a new avenue for applied research in discrete choice modeling.

choice modelingconsideration setconsistent estimation

This study addresses the characterization of identification sets in stochastic choice models—including random utility, bounded rationality, and dynamic discrete choice frameworks—with the aim of determining whether distinct distributions over choice rules are observationally equivalent. The authors propose a unified analytical framework based on elementary permutation transformations, integrating tools from probability distribution transformations, convex analysis, and global inversion theory. For the first time, they fully characterize the inequality representations and extreme-point structures of identification sets across multiple model classes. Furthermore, they develop a globally valid identification test under smoothly varying parameters by leveraging global inversion techniques. This work not only establishes a rigorous theoretical foundation but also delivers a practical tool for conducting global identification tests in empirical applications.

bounded rationalitydynamic discrete choiceidentification

This study addresses the limitation of traditional choice models in neglecting inter-temporal dependencies among customers’ historical transactions, which hinders effective utilization of preference information embedded in panel data. To overcome this, the authors introduce panel data into the Markov chain choice model for the first time, establishing a unified framework that integrates partial ranking preference information. They further propose a novel EM estimation algorithm that explicitly incorporates such partial orderings. Theoretical analysis reveals the computational complexity of conditional choice prediction and assortment optimization under this framework. Empirical evaluations demonstrate that the proposed algorithm consistently outperforms conventional Markov chain estimation methods and multinomial logit–based partial ranking benchmarks on both synthetic and real-world sushi datasets.

assortment optimizationchoice predictionMarkov chain choice model

Hot Scholars

SH

Stephane Hess

University of Leeds
choice modellingbehavioural modellingdiscrete choicestated preference
JY

Joseph Y. J. Chow

New York University
Behavioral informaticsurban transportation systems
MH

Min-hwan Oh

Seoul National University
Reinforcement LearningBandit AlgorithmsMachine Learning
BF

Bilal Farooq

Laboratory of Innovations in Transportation (LiTrans), Toronto Metropolitan University
SimulationBehavioural ModellingMachine LearningIntelligent Systems
NC

Ningyuan Chen

Department of Management, UTM & Rotman School of Management, University of Toronto
Revenue ManagementOnline LearningOperations ManagementBusiness Analytics