Score
Designs and implements models that learn and represent prototypical contact patterns—compact, interpretable templates of pairwise interactions—then map those prototypes to specific contact pairs in new inputs to explain model predictions. Work includes building prototype representations, prototype-to-instance mapping and attribution mechanisms, and evaluation procedures to measure how prototype-based interpretations reflect or influence modeled structures.
Existing prototypical networks predominantly rely on Euclidean-space prototypes, which constrain semantic interpretability and structural flexibility. Method: We systematically survey and compare Euclidean versus non-Euclidean prototype representation paradigms, introducing— for the first time—a unified analytical framework that elucidates how prototype geometry governs interpretability. Our approach integrates prototype learning, differentiable attention-based localization, multi-granularity part matching, and cross-dataset generalization evaluation. Contribution/Results: Experiments on three fine-grained benchmarks—CUB-200-2011, Stanford Cars, and Oxford Flowers—demonstrate that non-Euclidean prototypes substantially improve the trade-off between model interpretability and classification accuracy. Specifically, they enhance part-level semantic alignment and out-of-domain generalization robustness, offering greater structural expressivity and principled geometric grounding for prototype-based representation learning.
This work addresses the trade-off between data quality and quantity in robot learning, with particular emphasis on the critical role of contact information in dynamics and shape modeling. We propose a contact-aware Fisher information metric that incorporates pose and contact signals into the objective function to quantify the information content of individual data samples, enabling efficient data selection and refinement. In contrast to conventional paradigms relying on large-scale, low-quality datasets, our approach achieves substantial improvements in learning efficiency, stability, and generalization using only a small number of high-information samples. Our key contribution is the first explicit integration of contact awareness into the Fisher information framework, establishing a principled, interpretable, and computationally tractable criterion for evaluating data utility in contact-rich robotic learning—thereby advancing high-quality, data-driven embodied intelligence.
This paper systematically analyzes bottlenecks hindering prototype-based predictive models (PPMs) in explainable AI (XAI) from 2019–2024: insufficient prototype quality and diversity, weak cross-task generalizability, and lack of methodological standardization. Through systematic literature review, challenge attribution modeling, and technical evolution analysis, we first establish a comprehensive taxonomy of PPM challenges and propose a five-dimensional research roadmap covering model architecture, human-centered alignment, and evaluation paradigms. Key contributions include: (1) identifying the critical transition pathway from post-hoc explanation to *inherently interpretable* PPMs; (2) introducing a novel human-cognitive alignment and human-AI collaboration framework; (3) designing a unified, multi-faceted interpretability evaluation metric system; and (4) open-sourcing a structured literature repository covering 100+ works. This study delivers the first holistic development blueprint for inherently interpretable AI grounded in PPMs.
Prototype-based explanations often suffer from poor human interpretability due to insufficient focus on salient features. This paper addresses the “lack of focus” problem in prototype-driven explainable AI by proposing a novel method to identify semantically aligned key overlapping regions—termed *alike parts*—between an input instance and its nearest prototype. Our contributions are twofold: (1) We introduce the first prototype selection objective that explicitly incorporates feature attribution scores (e.g., SHAP or LIME) to enhance global prototype diversity; (2) We formally define and extract instance-prototype semantic alignment regions, leveraging a weighted feature overlap matching mechanism for precise localization. Extensive experiments across six benchmark datasets demonstrate significant improvements in human comprehension while maintaining classification accuracy—either stable or slightly improved—relative to baseline methods.
In high-risk visual tasks, ProtoPNet offers interpretability but suffers from an “interaction bottleneck”: model defect correction requires time-consuming retraining. This paper proposes Proto-RSet, the first framework to integrate the Rashomon set concept into prototype learning, enabling millisecond-scale interactive editing and debugging by non-expert users. Proto-RSet leverages Rashomon set sampling, constraint-based optimization, and differentiable prototype matching to rapidly generate multiple accurate and diverse ProtoPNet variants—while preserving performance (accuracy variation ≤ ±0.5%). We validate its efficacy on bias-mitigated bird recognition and clinical skin cancer diagnosis (debugging), with endorsement from domain experts. The core contribution is breaking the interaction bottleneck: Proto-RSet enables real-time, interpretable model correction without retraining.
To address the need for interpretable understanding of functional regions (e.g., graspable or pressable areas) in robotic autonomous manipulation and human–robot interaction, this paper introduces the first 3D point-cloud-based functional region detection method. It pioneers the integration of probabilistic prototype learning into this task. Built upon a PointNet++ backbone, the approach jointly learns functional region localization and human-interpretable explanations via probabilistic prototype matching, soft attention mechanisms, and local geometric encoding. Unlike black-box models, it achieves state-of-the-art accuracy on 3D-AffordanceNet (improving mAP by 1.2%), while simultaneously generating faithful, semantically grounded explanations: each predicted region is explicitly linked to a human-understandable training prototype (e.g., “similar to a canonical grasping prototype”). This work establishes a novel, trustworthy paradigm for explainable 3D functional reasoning.
This work investigates how vision foundation models can achieve genuine understanding of object affordances by jointly modeling geometric structure and interactive behavior. It identifies, for the first time, geometric perception and interaction perception as two composable fundamental components of affordance understanding. To this end, the authors propose a novel zero-shot fusion strategy that requires no additional training: part-level geometric prototypes are extracted using DINO, and then fused with verb-conditioned spatial attention maps generated by Flux. Experimental results demonstrate that this approach achieves performance comparable to weakly supervised methods under zero-shot settings, thereby validating the effectiveness and novelty of the proposed mechanism.
This study investigates whether the geometric structure of concepts in large language models is fixed by pretraining priors or dynamically shaped by context. Through representational similarity analysis, activation interventions, and cross-model comparisons (Gemma, Qwen), the work demonstrates for the first time that contextual instructions can deliberately construct arbitrary conceptual topologies—such as ring or tree structures—and causally dominate generation behavior in large models (e.g., Gemma-31B, Qwen-27B), with effect sizes ranging from 0.6 to 0.9 in similarity metrics. This influence is not merely an epiphenomenon of representation but a controllable driver of output. In contrast, smaller models fail to reliably exhibit this capability, highlighting a qualitative divergence in how model scale mediates contextual control over conceptual geometry.
This study addresses the challenge of organizing persistent information into reusable contexts within world models by proposing the SPRII training principle. Leveraging inter-interaction relationships as weak supervision signals, this method introduces Align and Cross components that employ contrastive learning and cross-trajectory prediction mechanisms without requiring numerical labels. These components guide the model to construct shared context representations encoding the system's persistent attributes. Experimental evaluations across 13 scenarios demonstrate that SPRII improves downstream task performance by over 10% on average and increases persistent attribute reading accuracy by more than 15%.
This study addresses the problem of "model collapse"—a degradation in performance arising when multiple models interactively learn from synthetic data generated by one another. By formalizing inter-model interactions as a directed graph, the work establishes, for the first time, necessary and sufficient conditions for model collapse in multi-model settings, thereby extending beyond prior analyses limited to single-model self-training. The theoretical framework integrates directed graph topology, finite-sample analysis of linear regression, and asymptotic theory of M-estimators to rigorously characterize the collapse mechanism. Extensive numerical experiments validate the theoretical findings and uncover an intrinsic relationship between the structure of the interaction graph and the extent of performance degradation across models.
本文提出一种生成图像模型的可视化界面设计空间,通过分解用户界面、可控模型对象及映射函数来分析不同交互技术,以促进新界面设计。