Institution profile

Squirrel Ai Learning

Industry researchasia · cn
Official website
Research library99linked papers
Opportunities0open roles
Selected work

Representative Papers

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

May 11, 2026arXiv.org

Large audio language models often lose critical acoustic evidence under severe noise, leading to unreliable responses. This work proposes EchoDistill, a novel noisy-to-clean self-distillation framework that leverages clean audio as privileged information to guide a teacher model. The student model is optimized through masked token distillation and task-gated consistency shaping, while the backbone network remains frozen to ensure zero additional inference overhead. Experimental results demonstrate that this approach improves average accuracy by 1.63% under strong noise conditions at −10 dB, with Qwen2.5-Omni achieving 62.94%. These findings confirm that EchoDistill effectively enhances robustness against acoustic degradation without incurring extra computational costs during inference.

3 citationsRead paper

MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection

Mar 23, 2025

Accurately identifying errors in K–12 multimodal math assignments—comprising both handwritten or typeset text and diagrams—remains challenging, as current multimodal large language models (MLLMs) lack the capability to jointly reason over image and text modalities and precisely localize and attribute solution-step errors. Method: This paper proposes the first three-stage mathematical agent hybrid framework: (1) image-text consistency verification, (2) visual-semantic parsing, and (3) cross-modal error integration analysis—explicitly modeling multimodal associations between problem statements and solution steps. The framework integrates vision understanding, symbolic logical reasoning, and pedagogically grounded constraints within a specialized collaborative agent architecture. Contribution/Results: Evaluated on real-world educational data, the framework improves step-level error detection accuracy by 5% and error-type classification accuracy by 3%. It has been deployed at scale across a platform serving over one million students, achieving 90% user satisfaction and substantially reducing manual review overhead.

1 citationsRead paper

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

Oct 06, 2026

This study addresses a critical blind spot in existing LLM backdoor defenses, which focus solely on input triggers while overlooking risks from model-generated content. We propose an answer-side backdoor attack in multi-turn dialogues, introducing the novel paradigm of "model-planted triggers." Specifically, a benign initial turn induces the model to output specific tokens that enter the conversation history; these self-generated tokens are subsequently exploited to bypass safety alignment when harmful queries follow. The stealthy implantation is achieved through data poisoning, representation-level analysis, and adversarial prompt engineering. Experiments across four LLMs demonstrate that a mere 5% poisoning rate yields near-perfect attack success while preserving general utility. These findings expose fundamental vulnerabilities in current input-sanitization defense frameworks.

0 citationsRead paper

World Requirement Model: Learning Requirement-Change Consequences from Typed Artifact Graphs

Oct 03, 2026

This study addresses the challenge of accurately predicting how requirement changes propagate to interdependent elements such as stakeholders, constraints, and tests. To this end, it proposes an artifact-addressable world model that constructs typed artifact graphs to encode engineering context. By integrating relation-aware attention with type propagation mechanisms and coupling them with world-decision representation learning for dynamic evolution, the approach achieves context-aware consequence scoring through a shared readout layer. Evaluated on synthetic cases, the method attains an impact prediction Mean Average Precision (MAP) of 0.724, representing a 16.2% improvement over baselines, while the lowest-quartile Average Precision increases by 32.4%. Demonstrating significant superiority across multiple metrics compared to competing systems, this work establishes a comprehensive evaluation paradigm for requirement world prediction.

0 citationsRead paper

FutureWorlds: Learning Robotic World Models from Alternative Futures

Oct 01, 2026

This study addresses key challenges in robotic world models, including the difficulty of translating surrogate predictions into effective learning signals, low candidate diversity, and cumbersome trajectory history maintenance. To this end, we propose a unified optimization framework that integrates a multimodal discrete autoregressive architecture with diverse beam search to generate candidate future scenarios balancing confidence and diversity. We introduce a candidate-specific bounded memory mechanism to ensure historical consistency between generation and scoring, alongside the MemSPO algorithm, which converts video rewards into group-relative advantages for policy optimization. Experimental results demonstrate that our approach significantly reduces LPIPS scores on benchmarks such as RT-1, substantially improving generation quality within only 200 update steps while achieving more precise motion prediction and consistent object states.

0 citationsRead paper
Recent publications

Latest Papers

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

Oct 06, 2026

This study addresses a critical blind spot in existing LLM backdoor defenses, which focus solely on input triggers while overlooking risks from model-generated content. We propose an answer-side backdoor attack in multi-turn dialogues, introducing the novel paradigm of "model-planted triggers." Specifically, a benign initial turn induces the model to output specific tokens that enter the conversation history; these self-generated tokens are subsequently exploited to bypass safety alignment when harmful queries follow. The stealthy implantation is achieved through data poisoning, representation-level analysis, and adversarial prompt engineering. Experiments across four LLMs demonstrate that a mere 5% poisoning rate yields near-perfect attack success while preserving general utility. These findings expose fundamental vulnerabilities in current input-sanitization defense frameworks.

0 citationsRead paper

World Requirement Model: Learning Requirement-Change Consequences from Typed Artifact Graphs

Oct 03, 2026

This study addresses the challenge of accurately predicting how requirement changes propagate to interdependent elements such as stakeholders, constraints, and tests. To this end, it proposes an artifact-addressable world model that constructs typed artifact graphs to encode engineering context. By integrating relation-aware attention with type propagation mechanisms and coupling them with world-decision representation learning for dynamic evolution, the approach achieves context-aware consequence scoring through a shared readout layer. Evaluated on synthetic cases, the method attains an impact prediction Mean Average Precision (MAP) of 0.724, representing a 16.2% improvement over baselines, while the lowest-quartile Average Precision increases by 32.4%. Demonstrating significant superiority across multiple metrics compared to competing systems, this work establishes a comprehensive evaluation paradigm for requirement world prediction.

0 citationsRead paper

FutureWorlds: Learning Robotic World Models from Alternative Futures

Oct 01, 2026

This study addresses key challenges in robotic world models, including the difficulty of translating surrogate predictions into effective learning signals, low candidate diversity, and cumbersome trajectory history maintenance. To this end, we propose a unified optimization framework that integrates a multimodal discrete autoregressive architecture with diverse beam search to generate candidate future scenarios balancing confidence and diversity. We introduce a candidate-specific bounded memory mechanism to ensure historical consistency between generation and scoring, alongside the MemSPO algorithm, which converts video rewards into group-relative advantages for policy optimization. Experimental results demonstrate that our approach significantly reduces LPIPS scores on benchmarks such as RT-1, substantially improving generation quality within only 200 update steps while achieving more precise motion prediction and consistent object states.

0 citationsRead paper

DrafTS: Time-Aware Decomposition with Residual Correction for Time Series Modeling

Sep 27, 2026

This study addresses the challenge of pattern degradation in time series caused by noise-dynamics coupling, where conventional filtering tends to suppress meaningful temporal dynamics. To overcome this, we propose a model-agnostic framework for time-aware decomposition and residual correction. By leveraging instantaneous amplitude and frequency characteristics to guide signal decomposition, the method isolates principal components for modeling while employing a lightweight module to correct residuals, thereby preserving evolutionary dynamics during denoising. This framework can be seamlessly integrated into diverse backbone architectures. Extensive experiments across six model types and four time series tasks demonstrate consistent and significant performance improvements, validating its strong generalizability and effectiveness.

0 citationsRead paper

RecToolBench: Benchmarking Recommendation-Specific Tool Orchestration under Fuzzy User Intent

Sep 24, 2026

This study addresses the limitation of existing recommender system benchmarks, which predominantly assume explicit user intents and simplified tool environments, thereby inadequately evaluating tool orchestration capabilities under ambiguous intents. To this end, this work proposes the first recommendation-specific tool orchestration benchmark tailored for ambiguous user intents. Built upon the Model Context Protocol (MCP), the benchmark constructs over a thousand tasks through a synthesis–obfuscation–evaluation pipeline, encompassing invocation patterns ranging from single-step to complex hybrid multi-step calls. Evaluation is conducted by combining rule-based verification with LLM-assisted scoring. Experimental results demonstrate that large language models exhibit significant bottlenecks in semantic parameter grounding and multi-step evidence integration, confirming that tool orchestration remains a core challenge for agentic recommendation systems.

0 citationsRead paper