🤖 AI Summary
This study addresses the time-consuming quantification of conditional probability tables in Bayesian networks for software decision-making and the lack of empirical comparisons among available methods. Within a software development organization, we systematically compare two semi-automated quantification approaches: the weighted sum algorithm (WSA) and the ranked nodes method (RNM). Employing expert-driven modeling and model walkthrough techniques, their performance is evaluated in feature selection and UI design decisions. Results indicate that although both methods yield similar rankings, their uncertainty distributions differ significantly. This work reveals the critical impact of method selection on uncertainty representation, demonstrating that shortlists alone are insufficient to establish model equivalence. It further emphasizes the necessity of calibrating complete output distributions in conjunction with node semantics.
📝 Abstract
Expert-driven Bayesian networks can support recurring software decisions when historical data are limited, but quantifying their conditional probability tables (CPTs) requires many probability judgments. Semi-automatic methods reduce direct elicitation, yet practitioners have little comparative evidence from operational decision processes. We report an embedded case study comparing the Weighted Sum Algorithm (WSA) and Ranked Nodes Method (RNM) as complete elicitation-and-quantification pipelines in two contexts within a single software research and development (R&D) organization: feature selection for Internet of Things projects and user interface design selection. Within each context, the pipelines shared the graph, root priors, and decision evidence. We evaluated them through 15 expert-defined model-walkthrough scenarios and retrospective reconstructions of alternatives recorded in decision meetings. WSA matched 7/7 and 6/8 walkthrough expectations, whereas RNM matched 4/7 and 4/8. Both pipelines placed the selected alternatives near the top of the retrospective rankings. Their complete distributions nevertheless differed: the median total variation distance was 0.304 across four distinct feature patterns and 0.340 across seven distinct design patterns, rising to 0.890 for one design pattern. These results concern two ordinal value-estimation models in one organization. They show that similar shortlist behavior does not imply equivalent representations of uncertainty. For comparable software decision-support models, method selection and validation should consider node semantics, the judgments experts can provide, calibration requirements, complete output distributions, and the intended downstream use of the probabilities.