Score
Designs and implements systems that generate and use content-based identifiers (e.g., cryptographic hashes) to address and verify data, enabling deduplication, integrity checks, and location-independent lookup. Builds algorithms and protocols to create those identifiers, compare write intents to detect conflicts, determine which operations can run in parallel, and reconcile concurrent writes.
Git commit hashes are commonly regarded as immutable, unique identifiers that underpin content integrity through cryptographic hashing. However, this work demonstrates that under signing mechanisms, such hashes can be manipulated without compromising signature validity, thereby undermining the foundational assumption of hash-based content addressing. Introducing the novel concept of “hash-chain malleability,” the authors exploit inherent flexibility in Git’s commit representation to generate syntactically and semantically valid commits with distinct hashes—despite unchanged logical content and valid signatures—without access to the private key or breaking SHA-2. Techniques include ECDSA signature negation, insertion of unhashed subpackets in OpenPGP, and non-canonical DER length re-encoding in S/MIME. All resulting commits pass `git verify-commit` and retain GitHub’s persistent “Verified” badge, demonstrating a practical end-to-end attack that challenges the security guarantees of Git’s hash-based integrity model.
To address high network overhead, excessive cloud-side computational burden, and elevated latency caused by secure deduplication in encrypted cloud storage, this paper proposes a source-end deduplication paradigm. It migrates duplicate detection and lightweight proof-of-ownership (PoW) computation to client-side edge trusted execution environments (TEEs), integrating content-defined encryption (CDE), distributed fingerprint indexing, and an edge-coordination protocol to establish a cloud-edge collaborative, locality-aware deduplication architecture. This work is the first to achieve PoW offloading within TEEs and secure ciphertext fingerprint comparison, thereby breaking the traditional security-performance trade-off inherent in source-end or server-side approaches. Experimental results demonstrate that, compared to state-of-the-art schemes, the proposed solution reduces upload traffic by 72%, decreases cloud-side computational overhead by 65%, cuts end-to-end latency by 58%, and resists ciphertext-deduplication inference and PoW side-channel attacks.
The widespread deployment of algorithms—particularly large language models—in high-stakes domains such as healthcare, criminal justice, and finance has intensified challenges surrounding accountability, transparency, and traceability. This work proposes a novel framework that systematically integrates Digital Object Identifiers (DOIs) into algorithmic governance by establishing a unique identity system for algorithms enriched with metadata parameters. Complemented by a dedicated cryptographic authentication protocol and secure API mechanisms, the framework enables end-to-end auditable tracking across the algorithm’s entire lifecycle. It thereby facilitates reliable provenance tracing, bias mitigation, and scientific reproducibility, while laying a verifiable and audit-ready governance foundation for AI agents and multimodal large language models.
In digital forensics, the atomicity and integrity of storage snapshots lack rigorous definitions that jointly guarantee both instantaneousness and causal ordering—undermining evidentiary admissibility in legal proceedings. To address this, we propose a novel atomicity definition grounded in causal consistency, overcoming the limitation of conventional time-based atomicity models. We further rectify conceptual flaws in existing integrity definitions and introduce a revised, theoretically sound yet engineering-practical integrity criterion—explicitly supporting copy-on-write (CoW) implementations. Our approach integrates causal modeling, formal snapshot semantics, CoW mechanism analysis, and formalization of forensic quality criteria, yielding a verifiable snapshot semantic framework. This work establishes the first theoretical foundation for forensic tool design that unifies causal ordering with instantaneous state capture, thereby significantly enhancing the forensic validity and judicial admissibility of live data acquisition.
To address risks of large language model (LLM) theft and misuse, this paper proposes a verifiable, tamper-resistant model fingerprinting technique. Methodologically, it constructs a cryptographic hash chain using question-answer pairs and SHA-256, enforcing fine-grained response constraints and hash binding to ensure strong integrity verification. It is the first work to formally define and satisfy five core fingerprint properties: transparency, efficiency, persistence, robustness, and unforgeability. Extensive experiments across multiple LLMs demonstrate that the fingerprint withstands benign modifications—including fine-tuning and pruning—as well as adversarial erasure attacks, while preserving near-original model performance post-embedding. This work delivers the first complete solution for LLM copyright protection and provenance tracking that simultaneously achieves theoretical rigor and engineering practicality.
Existing distributed systems lack a unified identifier scheme that simultaneously achieves storage efficiency, temporal orderability, embedded provenance metadata, query-free verifiability, external confidentiality, and cross-century addressability. This work proposes SKID, a three-layer identity architecture that, for the first time, integrates all six critical properties within a single framework. SKID employs a 64-bit time-topology integer primary key, a 128-bit BLAKE3-authenticated UUID extension, and an AES-256 single-block encryption layer to enable deterministic bidirectional conversion. The design ensures B-tree-friendly ordering and compact 8-byte storage while supporting zero-query verification and internal parseability without external readability, making it suitable for multi-century, cross-trust-boundary distributed systems.
This study addresses the lack of a stable identity mechanism for AI agent skills that preserves semantic similarity, a limitation exacerbated by conventional hashing methods' inability to support fine-grained comparison. To overcome this, the authors propose a locality-sensitive fingerprinting approach based on multiple SimHash instances, which decomposes each skill into three components—prompt, code, and tool—and embeds them separately to generate a fixed 120-byte tripartite fingerprint. This representation enables accurate recovery of skill family identity, precise attribution of reused components, and reliable detection of independently implemented variants, even after rewriting, renaming, refactoring, or cross-language reimplementation. Experimental results demonstrate an AUC of 0.974 across 4,950 skill pairs, a 77-fold reduction in bit overhead, and exact identification of modified components in all 906 tampered samples.
This work addresses a critical challenge in verifying compliance with future international AI agreements: preventing AI clusters from covertly leaking undisclosed computation results via I/O channels in the absence of mutually trusted processors. The paper proposes a novel security architecture that eliminates the need for shared trusted hardware between Prover and Verifier. By employing passive optical splitters to capture all I/O traffic in real time, the system generates tamper-evident data fingerprints through hash commitments and coin-tossing protocols. A secure gateway further neutralizes covert channels—including emulation, timing side channels, and protocol header steganography. This approach enables the first auditable yet privacy-preserving verification of AI cluster I/O, offering low cost, ease of deployment, and rapid prototyping by small teams within months, thereby effectively blocking covert exfiltration paths while supporting post-hoc compliance audits.
This work addresses the lack of formally verifiable, fine-grained access control mechanisms in local-first systems operating at scale under low-trust collaboration settings. We propose a bottom-up approach that integrates a capability-based authorization model with Hashed Chronicle—a replicated data type—to design a Byzantine fault-tolerant collaborative group management mechanism. For the first time, system-level formal verification is introduced into local-first access control by leveraging the Verus framework to specify and verify a Rust implementation with zero runtime overhead. We formalize the semantics and key invariants of a simplified CRDT and prove the correctness of the core authorization logic, thereby providing Matrix, Keyhive, and similar systems with an integrable, high-assurance security foundation.
This work addresses the interoperability challenges in digital credential ecosystems, which stem from heterogeneous standards and independent evolution, and which traditional approaches fail to fully explain—particularly regarding incompatibilities that persist even under shared data models and the precise trust requirements of verifiers. To resolve this, the paper proposes a verifier-centric conceptual model that decomposes credential verification into three layers: signature validation (L1), semantic interpretation (L2), and validity assessment (L3). It further introduces two orthogonal planes—institutional and logistical—to construct a five-function framework within a three-dimensional deployment space. Building on this foundation, the authors design the Shinken framework, which integrates trust declarations, verification material exchange, and deployment strategies to enable cross-stack analysis. Evaluations across four learner credential stacks and an accreditation federation demonstrate that the model effectively elucidates and mitigates key issues including interoperability barriers, verification overhead, privacy risks, and terminological ambiguity.