๐ค AI Summary
This study investigates the optimal timing for incorporating preference embeddings in artificial intelligence systemsโwhether during training or as a post-processing step. Drawing on information design theory, the work proposes a unified welfare framework that is agnostic to specific decision objectives, revealing how preference embedding induces a contraction in posterior means and thereby affects the value of information. By integrating convexity analysis of information value with models of cognitive constraints, the paper demonstrates that, in the absence of cognitive frictions, preference-agnostic training weakly dominates preference-embedded approaches. However, under human cognitive limitations, embedding preferences during training enhances performance through automatic threshold computation, offering theoretical justification for modular AI architectures.
๐ Abstract
Machine learning systems embed preferences either in training losses or through post-processing of calibrated predictions. Applying information design methods from Strack and Yang (2024), this paper provides decision problem agnostic conditions under which separation training preference free and applying preferences ex post is optimal. Unlike prior work that requires specifying downstream objectives, the welfare results here apply uniformly across decision problems. The key primitive is a diminishing-value-of-information condition: relative to a fixed (normalised) preference-free loss, preference embedding makes informativeness less valuable at the margin, inducing a mean-preserving contraction of learned posteriors. Because the value of information is convex in beliefs, preference-free training weakly dominates for any expected utility decision problem. This provides theoretical foundations for modular AI pipelines that learn calibrated probabilities and implement asymmetric costs through downstream decision rules. However, separation requires users to implement optimal decision rules. When cognitive constraints bind, as documented in human AI decision-making, preference embedding can dominate by automating threshold computation. These results provide design guidance: preserve optionality through post-processing when objectives may shift; embed preferences when decision-stage frictions dominate.