certificate matching

Design and implement methods and tools to locate and extract digital certificates and signing metadata from software artifacts (e.g., binaries and system images), compare certificate signatures or public keys to identify matches, and link certificates to recovered keys. Use these matches to attribute keys to specific applications or system images and to quantify which apps or devices are affected by shared, reused, or compromised certificates.

certificatematching

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$231K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of copyright infringement detection in automated software plagiarism identification, which is complicated by the diversity of digital artifacts. The authors systematically review the legal and technical landscape and propose a classification framework for detection challenges based on artifact types. Building upon this framework, they integrate multiple similarity detection paradigms—including fingerprinting, software birthmarks, and code embeddings—into a unified, open-source platform named Project Martial. The system enables cross-artifact-type code plagiarism detection and demonstrates, through real-world case studies, that combining complementary techniques significantly enhances both detection accuracy and applicability. Project Martial thus provides a reproducible tool to support both academic research and forensic practice in software copyright enforcement.

code similaritycopyright infringementdigital artifacts

Centralized package registries (e.g., PyPI, npm) strengthen security controls, yet their authority collapses at distribution boundaries—including mirrors, corporate proxies, repackaging, and air-gapped transfers—rendering them insufficient for source authentication, integrity assurance, and accountability. Method: This paper proposes a trust extension model tailored to modern software distribution, formally characterizing the necessity and adaptability requirements of cryptographic signing across mirrors, proxies, and offline environments; it evaluates the synergistic defensive efficacy of centralized registries and end-to-end signing through historical practice and trust boundary theory. Contribution/Results: We establish software signing as a foundational trust primitive that transcends registry-level governance, providing a verifiable, traceable, and auditable technical basis for cross-boundary trusted distribution—thereby enabling robust provenance verification, tamper-evident integrity, and enforceable accountability across heterogeneous deployment contexts.

Registry security alone cannot guarantee trust across distribution boundariesSigning provides essential defense layer for software supply chain assuranceSoftware signing ensures artifact integrity and verifies producer identity

This study addresses the limited robustness of existing software watermarking techniques in cross-platform binary programs, which hinders effective detection of code plagiarism. The authors propose a novel cross-platform watermarking method based on Ghidra’s P-code intermediate representation, which unifies binary program representations across diverse architectures. By integrating program feature extraction with similarity metrics such as the Simpson index, the approach achieves highly consistent plagiarism detection. The work presents the first empirical validation of watermark effectiveness in real-world cross-platform environments, uncovering a “dilution effect” caused by Windows library functions and demonstrating the superior discriminative power of the Simpson index under noisy conditions. Experiments spanning multiple CPU architectures and programming languages yield a correlation coefficient as high as 0.9846, strongly confirming the method’s cross-platform robustness and practical utility.

binary analysiscross-platformintermediate representation

This study addresses the lack of cross-modal unified evaluation and open-source compliance verification for soft binding between watermarking and fingerprinting in content provenance. Under a unified protocol, we benchmark image, audio, and video modalities using strictly auditable open-source models such as PixelSeal. By integrating content fingerprint retrieval, statistical calibration, and source-level bootstrap interval estimation, the proposed framework systematically evaluates perceptual quality, robustness, and false positive rates. This work presents the first cross-modal unified assessment of both technologies, identifying optimal open-source solutions, quantifying failure mode discrepancies under diverse attacks, and revealing the impact of platform color-space conversions on watermark bit stability.

benchmarkcontent provenancefingerprinting

Levels of Binary Equivalence for the Comparison of Binaries from Alternative Builds

Oct 11, 2024
JD
Jens Dietrich
🏛️ Victoria University of Wellington | Oracle Labs Australia

In software supply chain security, binaries built from identical source code on different platforms often exhibit bit-level discrepancies, rendering traditional byte-wise comparison ineffective for determining functional equivalence and detecting cross-build security risks. To address this, we propose a multi-level binary equivalence model and introduce the first clone-detection framework that jointly incorporates semantic- and behavioral-level equivalence reasoning—thereby overcoming the limitations of strict bitwise equality. Our approach integrates static analysis, bytecode parsing, and equivalence relation modeling to construct a verifiable, semi-synthetic benchmark and an automated equivalence decision system. Evaluated on 14,156 pairs of Java binaries, our method identifies bit-level differences in 26.49% of samples, yet accurately confirms their functional equivalence—demonstrating substantial improvements in trustworthiness and reliability of supply chain binary comparison.

Assessing if alternative builds confirm or reveal compromised binariesDefining equivalence levels for comparing binaries from different buildsEvaluating practical equivalence despite cryptographic hash differences

Latest Papers

What's happening recently
View more

This work addresses the interoperability challenges in digital credential ecosystems, which stem from heterogeneous standards and independent evolution, and which traditional approaches fail to fully explain—particularly regarding incompatibilities that persist even under shared data models and the precise trust requirements of verifiers. To resolve this, the paper proposes a verifier-centric conceptual model that decomposes credential verification into three layers: signature validation (L1), semantic interpretation (L2), and validity assessment (L3). It further introduces two orthogonal planes—institutional and logistical—to construct a five-function framework within a three-dimensional deployment space. Building on this foundation, the authors design the Shinken framework, which integrates trust declarations, verification material exchange, and deployment strategies to enable cross-stack analysis. Evaluations across four learner credential stacks and an accreditation federation demonstrate that the model effectively elucidates and mitigates key issues including interoperability barriers, verification overhead, privacy risks, and terminological ambiguity.

digital credentialecosysteminteroperability

研究通过三个月实验考察了Git提交签名的使用情况,发现尽管大多数参与者能成功签名,但在设置、多设备配置和仓库验证中遇到困难,且存在误解。

commit provenancedeveloper experienceGit commit signing

This work addresses the challenging problem of recovering original source code from stripped binary functions, a task where traditional decompilation typically yields only approximate pseudocode. The paper proposes a novel paradigm that replaces pseudocode generation with direct source code retrieval. By extracting anchors such as strings and constants from binaries, the method retrieves candidate functions from a source code corpus and constructs a multimodal representation incorporating assembly instructions, decompiled code, and metadata. A large language model (LLM) is then employed for semantic re-ranking of candidates. The approach integrates Ghidra-based static analysis with an inverted index system and introduces an iterative anchor refinement strategy. Evaluated on a high-quality tcpdump dataset, it achieves 95.2% instruction coverage, and attains 35.5% coverage on general-purpose GitHub repositories, demonstrating effectiveness in both ideal and noisy real-world scenarios.

binary functionsbinary-to-source matchingreverse engineering

Hot Scholars

QC

Quanwei Cai

University of Science and Technology of China
Applied CryptographyPET
JL

Jingqiang Lin

Professor, University of Science and Technology of China
system securitycryptography
AC

Abel C. H. Chen

Information & Communications Security Laboratory, Chunghwa Telecom Laboratories
Cellular NetworksIntelligent Transportation SystemPost-Quantum CryptographyHealthcare System
XC

Xiuzhen Cheng

School of Computer Science and Technology, Shandong University
BlockchainIoT SecurityEdge ComputingDistributed Computing
GX

Guanping Xiao

Associate Professor, Nanjing University of Aeronautics and Astronautics
Software ReliabilitySoftware AnalysisProgram AnalysisSoftware Evolution