🤖 AI Summary
This study addresses the vulnerability of language models to misleadingly optimistic assertions from CRM stakeholders during sales qualification, wherein incentive-misaligned statements are erroneously treated as objective evidence, thereby compromising decision-making. To investigate this, we propose a diagnostic framework that distinguishes persuasion effects from information deficits, employing bucket analysis, same-information controlled experiments, and computational step-length control for systematic evaluation. Our findings reveal that neither scaling nor enhanced reasoning mitigates such biases. Across seven mainstream models, false approval rates reach 87–97% under contradictory assertions, confirming that this failure mode is fundamentally rooted in persuasion rather than information insufficiency. This work offers a novel perspective on the reliability of large language models in high-stakes business applications.
📝 Abstract
Language-model agents increasingly answer questions over customer-relationship management (CRM) records, such as whether to qualify a sales lead. We identify a failure mode not addressed by a stronger model: when the context contains an assertion by a party with an incentive toward optimism - here the sales representative, a witness recorded in the CRM - the model treats the assertion as evidence and clears deals the company's own records deem unacceptable. Across 100 lead-qualification tasks from CRMArena-Pro, the representative asserts an acceptable timeline in every call and an acceptable budget in 76; on the 31 tasks where such an assertion contradicts the price list and installation policy, a model reading only the transcript clears the deal in 29 of 31 cases. The signature is consistent across seven models from four providers (misled on 87-97%); scale and explicit reasoning confer no resistance. Only 3 of 35 genuine failures involve no assertion: the failure is persuasion, not missing information. We contribute a diagnostic method rather than an architecture: (i) a bucket analysis that separates persuasion from information gaps, (ii) a same-information control showing that supplying the records to the model lowers strict accuracy from 41 to 18 while raising recall - precision collapses - and (iii) a compute-step control that holds extraction fixed and varies only who computes Budget and Timeline. The margin ranges from 42 points on an inexpensive model to 2-5 points on models that already compute correctly; on the strongest models the arms are within confidence intervals, so the pattern is a consistent direction and a soundness property, not a proved performance floor. We pre-specify a generalization test that returns a negative result, characterize the precondition (a policy exactly specified in the inputs), and release all evaluation artifacts.