birthmark similarity matching

Designs and implements matching systems that compute and combine similarity scores between birthmarks (feature-based representations of artifacts), using metrics such as cosine, Dice, Jaccard, Simpson, and edit-distance, to score and rank candidate sources and detect paraphrased or syntactically transformed clones.

birthmarksimilaritymatching

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of existing software birthmark techniques in project-level code reuse detection, which are often undermined by partial reuse and spurious similarities arising from small code modules. To overcome these challenges, the authors propose a project-level birthmark comparison framework based on symmetric aggregation of module-level similarities. The approach innovatively incorporates a module-size weighting mechanism to suppress noise from small modules and introduces a Top-k partial similarity strategy that focuses on highly similar module pairs. By integrating the resilience of birthmarks with a credibility assessment, the method demonstrates superior performance in detecting partial code reuse. Experimental evaluation on 35 open-source Java projects shows that the proposed technique significantly outperforms state-of-the-art methods, achieving both robustness and stability.

code plagiarismincidental similaritypartial reuse

This study addresses the challenge of copyright infringement detection in automated software plagiarism identification, which is complicated by the diversity of digital artifacts. The authors systematically review the legal and technical landscape and propose a classification framework for detection challenges based on artifact types. Building upon this framework, they integrate multiple similarity detection paradigms—including fingerprinting, software birthmarks, and code embeddings—into a unified, open-source platform named Project Martial. The system enables cross-artifact-type code plagiarism detection and demonstrates, through real-world case studies, that combining complementary techniques significantly enhances both detection accuracy and applicability. Project Martial thus provides a reproducible tool to support both academic research and forensic practice in software copyright enforcement.

code similaritycopyright infringementdigital artifacts

Improving Source Code Similarity Detection Through GraphCodeBERT and Integration of Additional Features

Aug 12, 2024
JM
Jorge Martinez-Gil
🏛️ Software Competence Center Hagenberg GmbH

To address the insufficient joint modeling of semantic and structural features in source code similarity detection, this paper proposes a multi-feature fusion method based on GraphCodeBERT. Building upon the pre-trained model, we introduce a learnable, task-specific output feature layer and employ a feature concatenation mechanism to end-to-end integrate newly extracted structural and semantic features with the original contextual representations. Unlike conventional fine-tuning paradigms, our design uniquely embeds auxiliary output features directly into the classification pipeline and jointly optimizes them. Extensive experiments on standard benchmarks—including Bench4BL and POJ-104—demonstrate significant improvements in accuracy, recall, and F1-score. These results validate the effectiveness of fine-grained, synergistic multi-feature representation for code similarity assessment and offer a novel paradigm for adapting pre-trained models to code analysis tasks.

Enhancing code similarity detection using GraphCodeBERTImproving precision and recall in code comparisonIntegrating extra features to boost model accuracy

Neural representational similarity measurement suffers from inconsistent nomenclature and implementation, severely impeding cross-study comparability and result reproducibility. To address this, we propose the first dynamically evolving naming standard framework that enables scalable, verifiable, and globally unique identifiers for similarity measures—overcoming the rigidity of static standards in rapidly advancing domains. We develop an open-source Python benchmarking platform integrating 14 mainstream toolkits (encompassing ~100 distinct measures) and provide unified formal modeling and explicit differentiation among 12+ variants of key methods (e.g., CKA). Leveraging a modular architecture and multi-package-compatible interfaces, our platform supports out-of-the-box standardized computation and evaluation. This work significantly enhances transparency, reproducibility, and efficiency of cross-study comparisons in representational similarity analysis, establishing foundational infrastructure for neural representation research.

Addressing naming and implementation inconsistencies across studiesProviding a framework for benchmarking and comparing similarity methodsStandardizing diverse similarity measures in evolving fields

In biomedical image segmentation validation, metrics such as the Hausdorff distance suffer from implementation inconsistencies across open-source toolkits, compromising benchmark reliability, introducing biomarker bias, and posing clinical deployment risks. To address this, we systematically evaluate 11 widely used toolkits and introduce, for the first time, a reference implementation based on high-fidelity 3D surface meshes. Our framework integrates real-world clinical data and a cross-platform consistency analysis. Statistical analysis reveals significant inter-tool variation in Hausdorff distance computations (p < 0.001), with interpolation strategy, boundary handling, and sampling density identified as primary sources of discrepancy. Based on these findings, we propose a reproducible and verifiable paradigm for distance-based evaluation, accompanied by standardized computational guidelines. This work substantially enhances the reliability, comparability, and clinical translatability of segmentation assessment.

Assess impact of metric discrepancies on medical segmentation validationIdentify inconsistencies in distance-based metric implementations across toolsProvide guidelines for selecting reliable open-source metric computation tools

Latest Papers

What's happening recently
View more

This study addresses the limited robustness of existing software watermarking techniques in cross-platform binary programs, which hinders effective detection of code plagiarism. The authors propose a novel cross-platform watermarking method based on Ghidra’s P-code intermediate representation, which unifies binary program representations across diverse architectures. By integrating program feature extraction with similarity metrics such as the Simpson index, the approach achieves highly consistent plagiarism detection. The work presents the first empirical validation of watermark effectiveness in real-world cross-platform environments, uncovering a “dilution effect” caused by Windows library functions and demonstrating the superior discriminative power of the Simpson index under noisy conditions. Experiments spanning multiple CPU architectures and programming languages yield a correlation coefficient as high as 0.9846, strongly confirming the method’s cross-platform robustness and practical utility.

binary analysiscross-platformintermediate representation

This study addresses the challenge of detecting large language model (LLM)-generated code that evades conventional plagiarism detection through semantics-preserving rewrites. It presents the first systematic evaluation of Java bytecode-based k-gram software watermarking (with k ranging from 1 to 6) in the context of LLM-generated code. The approach integrates five similarity metrics—cosine similarity, Dice coefficient, Jaccard index, Simpson index, and edit distance–based similarity—and is evaluated on code produced by three prominent LLMs. Results demonstrate that the proposed watermarking technique effectively identifies LLM-assisted plagiarism, with code generated by domain-specialized models (e.g., ChatGPT-5.1-Codex-Mini) exhibiting greater stealthiness, thereby confirming that model specialization enhances the concealment of plagiarized content.

code paraphrasingk-gramLLM-assisted plagiarism

This study addresses the longstanding reliance on subjective judgment in assessing traditional drawing skills by proposing a computer vision–based approach for quantitative evaluation. The work introduces a novel framework that systematically compares the performance of SIFT keypoint matching and Siamese neural networks in aligning hand-drawn sketches with reference templates. Experimental results demonstrate that SIFT significantly outperforms the Siamese network in capturing structural accuracy, thereby validating the feasibility of image-matching techniques for automated artistic skill assessment. This finding offers a promising pathway toward intelligent, objective tools for art education, bridging computational methods with creative skill evaluation.

art skills assessmentcomputer visiondrawing evaluation

Hot Scholars

JL

Jiayuan Li

wuhan uniersity
remote sensing, image processing, computer vision
XO

Xiaomin Ouyang

Department of Computer Science and Engineering, HKUST
Embedded AIAI for HealthIoTMobile Computing
CD

Christopher De Sa

Associate Professor of Computer Science, Cornell University
machine learning systems
YW

Yunchao Wei

Professor, Beijing Jiaotong University, UTS, UIUC, NUS
Computer VisionMachine Learning
PL

Percy Liang

Associate Professor of Computer Science, Stanford University
machine learningnatural language processing