Score
Designs and implements analyses and tools to recover, attribute, and measure the provenance of discrete tokens in a system, including statistical tests of separability of token origins and grounding theoretical provenance quantities in real tokenizers. Builds diagnostics that detect when untrusted or external tokens enter model control paths, expose control‑path entry points, and identify finite‑coverage invariance gaps in token provenance.
To address the challenge of tracing derivative relationships among large language models (LLMs), this paper introduces the first black-box model provenance verification framework. Unlike prior approaches, it requires no access to model weights or training data—only API-based output queries—and leverages statistical similarity of output distributions to perform model provenance inference. Crucially, it is the first to formalize this task as a multiple hypothesis testing problem, enabling high-confidence detection of derivative relationships. Evaluated on two real-world benchmarks encompassing over 600 models spanning 30M–4B parameters, the framework achieves 90–95% precision and 80–90% recall. Its core contribution is establishing a rigorous black-box provenance paradigm, supporting intellectual property protection, accountability for model misuse, and identification of foundational model issues—thereby providing a deployable technical foundation for LLM governance.
This work addresses the challenge of reliably tracing the provenance of tensors and operators through graph rewrites—particularly non-injective transformations—in AI compilers. The authors propose a lightweight, generative provenance method grounded in observational semantics, which infers origins by analyzing the behavioral effects of graph transformations rather than relying on identifier propagation. For the first time, they introduce coalgebraic modeling and bisimulation to this domain, guaranteeing provenance consistency even after intermediate nodes are eliminated. The approach requires no invasive compiler modifications and naturally supports non-injective rewrites. Evaluated within COVAN, a prototype AI compiler, the method demonstrates stable, low-overhead provenance tracking throughout an end-to-end compilation pipeline.
This work addresses the lack of provenance methods with provable error control for unauthorized use and multi-source attribution of large language models. It formally defines the model provenance problem with statistical guarantees and introduces the Model Provenance Set (MPS) framework, which constructs a small set of candidate source models satisfying a user-specified confidence level through sequential hypothesis testing and an adaptive exclusion mechanism. The proposed method provides the first provably correct coverage guarantee for model provenance, effectively handles multi-source scenarios, and achieves the target provenance coverage while strictly controlling the inclusion of irrelevant models. This approach is well-suited for model attribution and auditing tasks requiring rigorous statistical assurance.
This study addresses the overestimation of small models’ tool-calling capabilities caused by keyword-matching dependencies in existing benchmarks. We propose a low-cost diagnostic cascade integrating first-token probing, embedding drift detection, and factor analysis, revealing that specific training stages erase tool-calling priors. Targeted supervised fine-tuning (SFT) is subsequently applied for precise remediation. Experiments demonstrate that the model’s effective invocation rate increases from 0.1 to 0.959, while zero-shot pass rates significantly outperform baselines, confirming that the intervention preserves underlying representations. This work provides a systematic framework for evaluating and restoring the genuine tool-calling abilities of small language models under limited computational budgets.
Dafny provides strong correctness guarantees but imposes high verification overhead due to extensive manually written auxiliary assertions. This paper introduces DAISY, the first systematic framework for automatic assertion inference by large language models (LLMs) in scenarios with multiple missing assertions. Its core contributions are: (1) a hybrid fault localization method integrating LLM-generated candidates with Dafny’s error diagnostics; (2) a verification-oriented assertion taxonomy tailored to Dafny’s specification language; and (3) the empirical finding that verification is robust to assertion redundancy—full recovery of original assertions is unnecessary for successful proof. Evaluation shows DAISY achieves 63.4% verification success rate under single-missing-assertion conditions and 31.7% under multi-missing-assertion conditions. Notably, for some programs, DAISY completes equivalent verification with fewer assertions than the original, thereby reducing proof engineering effort.
This study addresses the vulnerability of synthetic data provenance to text paraphrasing and its limited utility in guiding data selection for recursive training. Focusing on financial text generation, we conduct comparative experiments integrating style rewriting with recursive model training to systematically evaluate how filtering strategies—based on source identifiers versus reference model scores—affect model degradation. Our findings reveal that data identity recognition and training value assessment constitute distinct problems, challenging the prevailing assumption that traceability inherently implies superiority. Experiments demonstrate that source attribution accuracy drops precipitously after rewriting and fails to consistently mitigate model collapse, establishing that provenance is not a sufficient condition for ensuring recursive training quality. This work thereby provides a new paradigm for synthetic data curation.
This study addresses the challenges of provenance detection and watermarking for visual content generated by large language models (LLMs) via two distinct pathways: direct generation and code-based rendering. It proposes a production-centric conceptual framework that organizes watermarking mechanisms according to production stages and introduces explicit verification specifications. Through interface documentation review, boundary case analysis, and multimodal theoretical synthesis, the work systematically compares detection discrepancies between these two pathways across images, videos, and code. The primary contribution lies in formulating a conceptual research agenda grounded in existing methodologies and interface analysis, thereby offering a systematic theoretical perspective and delineating future research directions for multimodal AI content provenance.
This study addresses the issue in data provenance verification where system outputs are frequently misclassified as observed values, thereby compromising downstream decision-making. To mitigate this, we propose structured defense mechanisms encompassing row-level hierarchical tagging, enforced single-entry write points, and query filtering, providing the first empirical quantification of their impact on production validation judgments. Experimental results demonstrate that these mechanisms increase classification coverage from 36.1% to 98.4% and successfully intercept 3,070 unauthorized writes. However, critical validation decisions remain unchanged, revealing a structural deficiency wherein defensive interventions do not intersect with the data informing those decisions. This finding substantiates the ineffectiveness of conventional provenance strategies in specific operational contexts.
本文提出SecTB-RTL框架,通过31项任务和124个硬件安全回归测试,审计AI生成的RTL验证计划的有效性,发现仅满足提供者模式并不保证执行有效性。
研究通过控制令牌注入方法,揭示了工具使用型语言模型的安全性是模型和其解码框架共同属性,并提出了解决方案如输入净化、解析器加固等。