🤖 AI Summary
Existing benchmarking methodologies rely on static datasets and struggle to support architectural trade-off analysis and evolutionary assessment of heterogeneous information systems in multi-model environments. This work proposes the TransforMMer framework, which reconceptualizes benchmark engineering as a systematic design tool by introducing a unified representation model that explicitly captures schema semantics and cross-model mappings. From a single source dataset, TransforMMer automatically generates semantically consistent yet structurally diverse variants across relational, document, and graph database models. The framework supports structural redesign operations—including embedding, augmentation, and hybrid partitioning—to enable reproducible cross-representation transformations. Experimental results demonstrate that query performance disparities primarily stem from interactions between workload characteristics and data representations, thereby validating the framework’s efficacy in guiding the evolution of heterogeneous systems.
📝 Abstract
Contemporary information systems operate in heterogeneous and continuously evolving data environments, where representation choices and structural redesign decisions strongly influence system behavior. Existing benchmarking approaches, however, rely mostly on static datasets and fixed schemas, providing limited support for analyzing architectural trade-offs or guiding evolution in multi-model settings.
This paper introduces TransforMMer, a framework for evolution-aware and representation-aware benchmark engineering in heterogeneous information systems. The approach treats benchmark construction as a systematic design process: starting from raw data, inferring structure, refining it conceptually, and generating comparable dataset variants across relational, document, and graph systems. The framework is grounded in a unified representation that enables explicit modeling of schemas and cross-model mappings and supports reproducible transformations across alternative representations.
We position benchmarking as a system-design tool for evaluating architectural and representation-level decisions in evolving information systems, rather than as a static comparison of database engines. Through controlled benchmark construction scenarios on real-world datasets, we demonstrate how structural redesign steps -- such as embedding, enrichment, and hybrid partitioning -- affect observed query costs across systems. The results show that performance differences emerge primarily from the interaction between workload and representation design.
By enabling systematic generation of structurally distinct yet semantically aligned dataset variants, the proposed approach connects conceptual data modeling with empirical system evaluation and supports reproducible, evolution-aware analysis of heterogeneous information systems.