ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文介绍了ETH-TraceBench,一个用于评估以太坊DeFi表示在时间、协议和合约变化下的基准,通过大规模事件流数据集及模型测试解决机器学习中的快捷方式问题。
📝 Abstract
Ethereum decentralized finance (DeFi) provides a public, time-stamped record of transaction-level event streams, but the same public symbols can create strong machine-learning shortcuts. We introduce ETH-TraceBench, a benchmark for evaluating Ethereum DeFi representations under temporal, protocol, pool/infrastructure, and symbolic shift. The raw event universe covers January 2021-December 2025 and contains 1.35 billion transactions with logs and 5.01 billion raw log rows. Model evaluation uses a fixed 911,267-instance supervised sample, training on 2021-2024, selecting models on 2025H1, and testing on 2025H2. Simple models perform strongly on the aggregate temporal test: TraceStats-GB reaches 0.953 macro-F1 and TopicEmitterHashMLP 0.959 on the canonical DEX test set. Performance drops sharply under protocol novelty, with macro-F1 of 0.794, 0.743, and 0.766 for TraceStats-GB, TopicEmitterTrace-SGD, and TopicEmitterHashMLP, while strict unseen-pool scores remain 0.927, 0.897, and 0.935. Uniswap v4 and Ekubo v1, both absent from supervised training, are materially harder than the full test. Jointly masking emitter and topic identity reduces DEX macro-F1 to 0.916 and liquidation macro-F1 to 0.774 for TopicEmitterTrace-SGD. A standard Transformer over log-index-ordered events provides no consistent advantage over a deterministic shuffle of the same events, indicating that high aggregate scores can arise without sophisticated chronological modeling. A natural-prevalence audit estimates 2025H2 DEX prevalence among logged Ethereum transactions at about 22.5%, and a deterministic 400-transaction audit finds complete agreement with task label sources and independently re-queried raw-log counts. ETH-TraceBench therefore treats difficult transfer and controlled-input conditions, rather than a single aggregate score, as the main evaluation target.
Problem

Research questions and friction points this paper is trying to address.

Ethereum DeFi
Temporal Shift
Protocol Shift
Contract Shift
Machine Learning Shortcuts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Ethereum DeFi
Benchmark Evaluation
Temporal Shift
Protocol Novelty
Model Performance
🔎 Similar Papers
2024-10-04arXiv.orgCitations: 0
K
Kemal Kirtac
Department of Computer Science, University College London; Department of Computer Science, University of Warwick
Carsten Maple
Carsten Maple
Professor of Cyber Systems Engineering, University of Warwick
SecurityPrivacy and Trust