🤖 AI Summary
This study addresses the lack of tools in digital humanities research for close reading and comparative analysis of large language model (LLM) outputs. The authors present a browser-based comparative close reading workbench that, for the first time, incorporates token-level log-probability data into critical generative AI research from humanities and social science perspectives. Proposing to treat generated text as an object situated within a probability distribution, the work integrates counterfactual visualization techniques and combines Hyland’s metadiscourse analysis, syntactic parsing, and multiple probability visualization methods—including heatmaps, entropy sparklines, and 3D probability landscapes. The system offers six analytical modes that render the uncertainty and diversity of model outputs legible at the token level, establishing a novel paradigm for humanists to interrogate the generative mechanisms of large language models.
📝 Abstract
LLMbench is a browser-based workbench for the comparative close reading of large language model (LLM) outputs. Where existing tools for LLM comparison, such as Google PAIR's LLM Comparator are engineered for quantitative evaluation and user-rating metrics, LLMbench is oriented towards the hermeneutic practices of the digital humanities. Two model responses to the same prompt are side by side in annotatable panels with four analytical overlays (Probabilities for token-level log-probability inspection, Differences for word-level diff across the two panels, Tone for Hyland-style metadiscourse analysis, and Structure for sentence-level parsing with discourse connective highlighting), alongside five analytical modes, Stochastic Variation, Temperature Gradient, Prompt Sensitivity, Token Probabilities, and Cross-Model Divergence, that make the probabilistic structure of generated text legible at the token level. The tool treats the generated text as a research object in its own right from a probability distribution, a text that could have been otherwise, and provides visualisations including continuous heatmaps, entropy sparklines, pixel maps, and three-dimensional probability terrains, that show the counterfactual history from which each word emerged. This paper describes the tool's architecture, its six modes, and its design rationale, and argues that log-probability data, currently underused in humanistic and social-scientific readings of AI, is an important resource for a critical studies of generative AI models.