Score
Designing interactive systems where multiple coordinated visual representations (maps, globes, timelines, attribute panels) remain synchronized through overlays, metadata linkage, and shared controls to support exploratory analysis and human–AI collaboration.
This paper addresses the ambiguous human–AI collaboration relationships and insufficient exploration of interaction paradigms in hybrid active visual analytics systems. Adopting a qualitative approach integrating bibliometric analysis and thematic coding, it synthesizes two decades of literature to construct a novel three-dimensional classification framework—spanning collaborative objectives, automation levels, and human roles—and proposes the first consensus-based operational definition of hybrid active visual analytics. The study identifies critical limitations: conceptual inconsistency, narrow interaction patterns, and imbalanced research distribution across domains and time periods. It further provides the first systematic characterization of evolutionary trajectories of real-world practices across eras and application scenarios. The findings establish a theoretical foundation, design guidelines, and a developmental roadmap for advancing human–AI collaborative visual analytics, thereby filling a foundational gap in modeling collaboration paradigms within hybrid active systems.
This study systematically identifies and organizes sixteen core challenges surrounding visualization in synchronous remote collaboration. Drawing on insights from twenty-nine international experts, it focuses on five collaborative scenarios—exploratory data analysis, ideation, visualization presentation, data-driven decision-making, and real-time monitoring—and, for the first time, categorizes these challenges into four research and development dimensions: technology selection, social factors, AI assistance, and evaluation. Integrating emerging trends in extended reality (XR) and artificial intelligence (AI), the work proposes a structured framework that offers both theoretical grounding and practical guidance for future research and system design in multimodal, visualization-supported collaborative environments.
This study investigates the intertextuality mechanisms between human-authored textual intent and AI-generated images in collaborative visual storytelling involving novice users and vision-language models (VLMs). Using GPT-4o’s image generation capability, we conducted a three-phase qualitative study integrated with fuzzy-set qualitative comparative analysis (fsQCA) to identify three core collaborative strategies: prompt iteration, semantic expansion, and multimodal complementarity. We propose a theoretical framework of “text–image intertextuality,” characterizing four collaborative patterns and three empirically derived pathways to successful collaboration—namely, the Educational Collaborator, Technical Expert, and Visual Thinker. Findings demonstrate that AI-induced semantic overflow positively enhances creative ideation, while revealing critical challenges: insufficient cultural representation, weak visual consistency, and difficulties in narrative translation. The work provides empirical grounding and interface-design implications for developing human-centered, role-adaptive AI assistants in creative authoring contexts.
This work addresses the challenge of deeply integrating Human-Centered Artificial Intelligence (HCAI) with visualization and visual analytics. We propose the first two-dimensional design space that systematically maps HCAI’s four core capabilities—amplify, augment, empower, and elevate—to the four stages of visual cognition: observe, explore, model, and report. Our framework guides bidirectional integration of AI and visualization, establishing two novel paradigms: “visualization for AI explainability” and “AI-enhanced human–machine collaboration.” Methodologically, we synthesize generative AI, large language models, foundation models, and interactive visualization techniques, while emphasizing human-centered evaluation, ethical alignment, and traceable design. The outcome is an actionable research and development roadmap that enables the visualization community to systematically incorporate HCAI principles—and positions visualization as a core enabling technology for enhancing HCAI’s trustworthiness and usability.
Traditional data management systems fail to simultaneously satisfy the low-latency, strong consistency, and high availability requirements of Human-Data Interaction (HDI), primarily because interaction bottlenecks are driven by user availability—not query semantics—and architectural separation between interfaces and backend systems precludes joint optimization. Method: We propose a novel “system-interaction co-design” paradigm that integrates database theory with visualization and interaction modeling, thereby unifying previously siloed layers; we design and implement an HDI infrastructure enabling sub-second response times, end-to-end consistency guarantees, and real-time interactive feedback. Contribution/Results: Evaluated across multiple generations of prototype systems, our approach significantly improves reliability and efficiency of interactive AI applications. It establishes a scalable, foundational architecture paradigm for next-generation human-AI collaborative intelligent systems.
Existing visualization research predominantly focuses on *how to use* interactive features, neglecting the critical question of *how to construct* them. Method: We propose the first three-layer decoupled interaction authoring task model—intent–technique–component—derived from empirical coding and abstraction of 592 interaction units across 47 real-world applications. Contribution/Results: This model provides descriptive, evaluative, and generative capabilities, enabling the first unified formalization of interaction authoring intent, technical implementation, and component instantiation. It yields a reusable, theory-grounded classification framework that supports critical evaluation of existing visualization tools and informs the design and validation of next-generation low-code interaction authoring systems.
This study addresses core challenges in human–data interaction in the AI era, including perceptual latency, limited scalability, outdated interaction paradigms, and insufficient reliability and interpretability of generative outputs. It systematically examines the impact of AI technologies on human–data interaction, exploration, and visualization, redefining the role of humans in intelligent analysis. Integrating insights from cognitive science, perceptual theory, and interaction design principles, the work proposes a human-centered interaction framework that synergizes large language models (LLMs), vision-language models (VLMs), and multimodal visualization techniques. Moving beyond traditional evaluation metrics focused primarily on efficiency and scalability, this research advances a cognition-driven analytical paradigm for human–AI collaboration and articulates design principles and future research directions for human-centric intelligent analytics systems in the AI era.
This work addresses the challenge that current AI agents struggle to interpret users’ concurrent interaction intents on shared artifacts, thereby limiting dynamic co-creation. To overcome this, we propose CLEO—a collaborative intelligent agent grounded in mixed-initiative interaction principles—that dynamically switches among delegation, guidance, and collaboration modes by recognizing user concurrent behaviors in real time. We introduce the first collaborative model capable of real-time intent interpretation, identifying five behavioral patterns, six triggering mechanisms, and four enabling factors, and implement a decision framework comprising six interactive loops. Based on 214 rounds of interactions with professional designers, we quantitatively analyze mode usage (70.1% delegation, 28.5% guidance, 31.8% collaboration) and release design guidelines alongside a labeled dataset to support future research.
Existing multimodal human-AI interaction systems often treat alignment, explainability, and user agency in isolation, leading to poor user understanding of AI intent and diminished trust and sense of control. This work proposes a unified collaborative framework that, for the first time, co-designs multimodal alignment, real-time multimodal explainable feedback (encompassing visual, textual, and spoken modalities), and user intervention mechanisms within a continuous interaction paradigm. By establishing a closed-loop architecture for multimodal intent recognition and responsive feedback, the framework significantly enhances users’ comprehension of system behavior, perceived control, and overall transparency. Empirical validation in two high-stakes, time-sensitive scenarios—collaborative design and warehouse robotics—demonstrates the efficacy of this approach in fostering effective and trustworthy human-AI collaboration.
This study addresses the challenges knowledge workers face in efficiently integrating information across multiple documents and constructing structured knowledge under high cognitive load. To this end, the authors propose an interactive visual system that enables users and large language models (LLMs) to collaboratively and iteratively build dynamic knowledge graphs within a shared visual environment, synergistically combining intelligent retrieval with human-driven organization. The system integrates LLM-powered document understanding, dynamic graph generation, an interactive interface, and a human-AI co-editing mechanism. A user study (N=12) demonstrates that, compared to a pure retrieval baseline, the proposed approach significantly improves the quality and coverage of organized knowledge while effectively reducing users’ cognitive load.