🤖 AI Summary
This study addresses the lag in human evaluation tooling for generative models and the absence of open-source, self-hosted platforms supporting controlled experiments by developing a multimodal web-based evaluation system. The platform employs a browser-side editing and distribution architecture, incorporating built-in significance testing and Bradley-Terry scoring algorithms. Furthermore, it integrates GDPR-compliant data auditing and access control modules to support complex experimental logic and machine-readable specification export. By releasing the complete codebase as open source, this work provides standardized infrastructure that enables researchers to efficiently conduct privacy-compliant, controlled human evaluation experiments.
📝 Abstract
Human judgement is the reference measure for evaluating generative models, yet the software used to collect it lags behing the methodology. Researchers adapt listening-test frameworks designed for perceptual protocols such as MUSHRA, rely on closed commercial survey platforms, or implement single-use web applications. Live arenas such as Chatbot Arena and Music Arena rank publicly deployed systems at scale, but do not support controlled comparisons of a laboratory's own models with its own participants. We present PANEL, an open-source, self-hosted platform for such studies. A study is authored in the browser and distributed as a single link, with audio, video, image, and text stimuli, seven question types, and screening and skip logic. The platform reports per-question summaries, across-condition significance tests, pairwise win rates and Bradley--Terry scores, and supports power analysis from pilot data. Consent versioning, self-service withdrawal, retention enforcement, and audit logging support GDPR-compliant operation. Each study exports as a machine-readable specification. PANEL is available at https://github.com/matteospanio/panel.