Target-confidence Recourse Using tSeTlin machines: TRUST
This work addresses a critical limitation of traditional counterfactual explanations, which focus solely on flipping prediction labels while neglecting model confidence and robustness—rendering them vulnerable to perturbations in high-stakes scenarios. To overcome this, the authors propose TRUST, a novel framework that explicitly incorporates user-specified target confidence directly into the counterfactual generation process. By leveraging the interpretable clause structure of Probabilistic Tsetlin Machines (PTMs) and integrating Bayesian optimization, TRUST jointly minimizes input perturbation cost while optimizing both prediction confidence and robustness. The method achieves this by explicitly linking confidence to the stability of rule activation. Empirical results demonstrate that TRUST consistently yields highly robust counterfactuals with low recourse costs across multiple datasets; for instance, on the Haberman dataset, it attains a confidence level of 0.92 with an L2 distance of merely 0.10.