Score
Designs and implements representations and processing pipelines that extract and classify user intents from natural language or other inputs into explicit semantic schemas or graph-based representations. Builds translators and validators that convert those intents into executable automation actions or edit plans, decompose them into operation steps, coordinate intent mappings across applications, monitor intent lifecycle events, and detect adversarial or malicious intent.
Resource-constrained edge devices face challenges in accurately understanding user intent from UI interaction traces, while simultaneously ensuring privacy preservation and real-time responsiveness. Method: This paper proposes a two-stage decomposed architecture: (1) generating structured sequential summaries of interaction behaviors, followed by (2) lightweight intent inference based on these summaries. The approach integrates context-aggregated enhancement and task-adaptive fine-tuning to strengthen semantic modeling capabilities of small models. Contribution/Results: Experimental results demonstrate that, under identical privacy guarantees and low-latency constraints, the proposed method achieves higher intent recognition accuracy than state-of-the-art large multimodal language models. It establishes an efficient, privacy-aware, and real-time interaction understanding paradigm for on-device intelligent agents.
To address the insufficient quality and reliability of LLM-generated workflows for complex, multi-intent user queries, this paper proposes Opus, a prompt-based intent framework that introduces a reproducible and customizable intent-capture layer between natural-language queries and workflow generation. Opus decomposes mixed intents into structured intent objects through integrated signal extraction, structured parsing, and intent-driven generation. Its core innovations include formal definitions of “workflow signals” and “structured intents,” along with a lightweight intent abstraction mechanism. Evaluated on a benchmark of 1,000 synthetically generated multi-intent queries, Opus significantly improves the logical coherence, semantic consistency, and semantic similarity of generated workflows—particularly under high-complexity conditions.
This work addresses the challenges in human-AI collaborative data analysis, where rapidly evolving analytical processes often undermine shared understanding, leading to undocumented assumptions, misaligned intentions, and context-poor prompts. To mitigate these issues, the authors propose a rule-based coordination layer that explicitly models user intent as editable, structured rules and validates their consistency with shared intent in real time during prompt formulation. By integrating rule-based reasoning, notebook parsing, and structured intent representation, the approach externalizes user intentions and provides early warnings of potential conflicts. A user study demonstrates that the system significantly enhances analysts’ awareness of their collaborators’ intentions and encourages reflection on analytical strategies, offering a novel design paradigm for human-AI collaborative data analysis.
In GUI-less systems, user intents expressed in natural language cannot be directly executed, posing a fundamental challenge for human-computer interaction. Method: We propose a novel “intent → code → execution” paradigm that leverages large language models (LLMs) to generate executable workflow code end-to-end from natural-language intents (e.g., “send the vehicle registration certificate to the insurance company”). Using GPT-4o-mini, lightweight OS-level APIs (without GUI dependencies), and customized prompt engineering, we systematically realize and validate this approach. Contribution/Results: To our knowledge, this is the first work to empirically demonstrate LLMs’ capability to generate complete, syntactically correct, and functionally executable workflow code for GUI-less environments. Extensive multi-intent evaluation shows high intent-to-code generation success rates and correct execution rates, confirming broad applicability. GPT-4o-mini exhibits strong intent comprehension and structured workflow modeling ability. Our approach overcomes traditional application-layer interaction bottlenecks and establishes a new foundation for natural-language-driven human-AI collaboration.
The increasing complexity of smart manufacturing environments demands interfaces that can translate high-level human intents into machine-executable actions. This paper presents a unified framework that integrates instruction-tuned Large Language Models (LLMs) with ontology-aligned Knowledge Graphs (KGs) to enable intent-driven interaction in Manufacturing-as-a-Service (MaaS) ecosystems. We fine-tune Mistral-7B-Instruct-V02 on a domain-specific dataset, enabling the translation of natural language intents into structured JSON requirement models. These models are semantically mapped to a Neo4j-based knowledge graph grounded in the ISA-95 standard, ensuring operational alignment with manufacturing processes, resources, and constraints. Our experimental results demonstrate significant performance gains over zero-shot and 3-shots baselines, achieving 89.33\% exact match accuracy and 97.27\% overall accuracy. This work lays the foundation for scalable, explainable, and adaptive human-machine
This work addresses the challenges of constructing and managing generative AI agent systems for long-horizon, stateful, multi-step business processes by proposing a graph-structured workflow design methodology. Leveraging the LangGraph framework, it explicitly models core mechanisms such as state management, conditional routing, and human-in-the-loop interventions. The approach is instantiated in three representative applications: SQL analysis with repair loops, retrieval-augmented generation gated by evidential validation, and human-AI collaborative policy review supporting interruption and checkpoint-based recovery. By treating behaviors like routing, pausing, and audit trails as explicit product features rather than implicit prompt logic, this study not only delineates the applicability boundaries of LangGraph in high-complexity workflows but also substantially enhances system controllability, reliability, and auditability in real-world operational settings, establishing a reusable engineering paradigm.
Automatically constructing high-quality, reusable skills from heterogeneous, fragmented interaction traces—often missing critical security behaviors—is highly challenging. This work proposes the W2S framework, which introduces a novel intermediate representation called RWSA to decouple skills into workflow structure, execution semantics, and runtime attachments, thereby enabling task decomposition, control-flow modeling, verification, rollback, and state management. W2S achieves efficient skill construction through trajectory segmentation, local skill draft generation, structural alignment, branch fusion, redundancy compression, and confidence-aware retention. Experimental evaluation across 70 skills demonstrates that W2S improves behavioral replay consistency by 10.5% compared to baseline approaches based on summarization and prompting.
This work addresses the “intent gap” between user expectations and program behavior in AI-generated code by proposing intent formalization as a central pathway to transform informal requirements into verifiable formal specifications. We systematically identify intent formalization as a critical challenge for reliable coding in the AI era and introduce an end-to-end verifiable coding framework that integrates formal methods, test-driven development, AI-generated postconditions, domain-specific languages, and human-AI collaboration. The framework supports a spectrum of approaches ranging from lightweight testing to fully automated correctness-preserving synthesis. Preliminary experiments demonstrate that interactive, test-driven formalization effectively enhances program correctness, that AI-generated postconditions can uncover real-world bugs, and that provably correct code can be automatically synthesized from informal specifications.
This work addresses the challenges of translating natural language intents into network policies, which often leads to errors and conflicts—particularly in multi-intent scenarios—where fault diagnosis is difficult and assurance mechanisms are reactive. To overcome these limitations, the paper proposes an end-to-end closed-loop intent-based networking system that leverages large language models to reliably map high-level intents to executable policies. The system incorporates structured validation and conflict-aware activation mechanisms to ensure policy consistency. Moreover, it enables proactive multi-intent fault prediction and root-cause disambiguation, transforming network assurance from passive response to active, interpretable early warning. Experimental results demonstrate that the approach significantly enhances the trustworthiness of automated operations by delivering actionable early alerts, explainable fault analyses, and quantifiable lead time for remediation.