Score
Designs, builds, and analyzes evaluation frameworks and experiments that measure whether action possibilities (affordances) learned for known objects or tools transfer to novel objects, substitutes, or unconventional interaction scenarios. This work includes creating benchmark tasks and scenarios, defining success metrics and scoring protocols for tool-substitution and nonstandard interactions, and testing generalization beyond the training tool set.
Recent advances in large language models have led to strong performance on reasoning and environment-interaction tasks, yet their ability for creative problem-solving remains underexplored. We study this capability through the lens of creative tool use, where a model repurposes available objects by reasoning about their affordances and attributes rather than relying on canonical usage. As a first step, we introduce CreativityBench, a benchmark for evaluating affordance-based creativity in LLMs. To this end, we build a large-scale affordance knowledge base (KB) with 4K entities and 150K+ affordance annotations, explicitly linking objects, parts, attributes, and actionable uses. Building on this KB, we generate 14K grounded tasks that require identifying non-obvious yet physically plausible solutions under constraints. Evaluations across 10 state-of-the-art LLMs, including closed and open-source models, show that models can often select a plausible object, but fail to identify the correct parts, their affordances, and the underlying physical mechanism needed to solve the task, leading to a significant drop in performance. Furthermore, improvements from model scaling quickly saturate, strong general reasoning does not reliably translate to creative affordance discovery, and common inference-time strategies such as Chain-of-Thought yield limited gains. These results suggest that creative tool use remains a major challenge for current models, and that CreativityBench provides a useful testbed for studying this missing dimension of intelligence, with potential implications for planning and reasoning modules in future agents.
This study addresses the cognitive inefficiency of traditional normal distribution visualizations in probability comparison tasks and the lack of a systematic account linking design choices to user cognition. For the first time, it systematically integrates affordance theory from psychology into the design of static probabilistic visualizations. By analyzing the affordance characteristics of existing normal density plots, the authors propose a novel visualization form—the Croissant Chart. Combining cognitive psychology theory, visualization design principles, and a preregistered user study (N = 808), they demonstrate that this chart significantly improves both accuracy and response efficiency in probability comparison tasks. The work establishes an affordance-driven design methodology capable of predictably enhancing task performance.
Existing visualization research lacks a systematic cognitive-level framework incorporating affordance theory, hindering explanation of how design choices and reader characteristics jointly shape the hierarchical structure of information communication. Method: This paper introduces the first theoretical framework of *visual cognitive affordance*, integrating insights from psychology, human-computer interaction, and visualization. It formally defines core constructs—perceptual, interpretive, and action affordances—and employs interdisciplinary theoretical synthesis and modeling to derive actionable design evaluation principles. Contribution/Results: The framework is empirically validated through representative visualization case studies, demonstrating its efficacy in optimizing information hierarchy representation and enhancing user comprehension efficiency. It provides both theoretical grounding and practical guidance for visualization design, thereby filling a critical formalization gap in applying affordance theory to the cognitive dimension of visualization.
Existing visualization research predominantly focuses on *how to use* interactive features, neglecting the critical question of *how to construct* them. Method: We propose the first three-layer decoupled interaction authoring task model—intent–technique–component—derived from empirical coding and abstraction of 592 interaction units across 47 real-world applications. Contribution/Results: This model provides descriptive, evaluative, and generative capabilities, enabling the first unified formalization of interaction authoring intent, technical implementation, and component instantiation. It yields a reusable, theory-grounded classification framework that supports critical evaluation of existing visualization tools and informs the design and validation of next-generation low-code interaction authoring systems.
This work addresses the limitations of current computer-using agents in reliably executing long-tail, complex, and infrequent human-computer interaction tasks, primarily due to the scarcity of multimodal interaction data. To bridge this gap, the authors introduce CUActSpot—the first comprehensive evaluation benchmark encompassing five modalities: graphical user interfaces (GUIs), text, tables, canvases, and natural images—alongside diverse operational actions. They further develop a renderer-based, extensible synthetic data generation framework that automatically produces training samples annotated with natural language instructions and precise action trajectories. Leveraging this data, the Phi-Ground-Any-4B model demonstrates substantial performance gains over all open-source models with fewer than 32 billion parameters on complex interactive tasks, significantly enhancing agent reliability in understanding and executing long-tail operations.
This work proposes the first open-world general-purpose affordance foundation model, unifying the core challenges of "where to interact" and "how to interact." Given only a single RGB-D image and a language instruction, the model predicts task-relevant functional region masks and 3D post-contact motion trajectories. A large-scale, standardized data pipeline integrates robotic manipulation, human demonstrations, simulation, and real-world scan data to construct a unified language–mask–3D motion affordance representation. Evaluated across eight benchmarks, the model substantially outperforms existing methods, achieving average gains of 23.9 in gIoU and 26.3 in cIoU, improving contact-point hit rates by 12.7–61.3%, and demonstrating superior 3D motion prediction. Notably, it enables zero-shot generalization across objects, tasks, and scenes and can be deployed directly on real robots without fine-tuning.
This work addresses a critical gap in the evaluation of large language models (LLMs) as tool users by shifting focus from mere task success rates to the underlying cognitive mechanisms of tool discovery. Inspired by cognitive science, it proposes the first systematic framework that decomposes tool discovery into three quantifiable dimensions: curiosity, recognition capability, and usage efficiency. Through empirical analysis in both combinatorial game settings and realistic simulated environments such as Voyager, the study uncovers a counterintuitive inverse scaling law: larger models exhibit diminished recognition capability despite their increased scale. The proposed framework not only offers a more fine-grained lens for evaluating LLMs’ tool-use proficiency but also demonstrates robust validity and generalizability across diverse experimental settings.
This work addresses the limited ability of robots to understand and transfer the causal properties of non-standard tools. It proposes the first framework integrating counterfactual reasoning with disentangled causal feature learning, leveraging dynamics simulation to uncover tool–task causal relationships. Guided by a vision-language model and augmented with geometric–physical perturbations, the approach generates novel tool representations that enable creative tool identification and skill generalization across objects. By jointly incorporating causal discovery, vision-language model feature extraction, counterfactual generation, and keypoint matching, the method significantly outperforms baseline approaches in tasks such as reaching, scooping, and high-reach retrieval, thereby enhancing both the reliability of tool selection and the accuracy of cross-task transfer.