Untrusted Authors, Trusted Answers: A Calculus of Fidelity-Graded Translations

📅 2026-07-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of traditional translation validation methods, which treat individual translations in isolation and struggle to assess trustworthiness within heterogeneous, multilingual, multi-path translation graphs. The authors propose a fidelity-graded translation calculus that models translations as composable, verifiable graph structures, dynamically establishing trust levels through inline validation and multi-path consistency checks, thereby enabling end-to-end self-certifying answers. Built upon a Lean 4 mechanized core, the hurdy-gurdy platform integrates 13 languages and 13 translator pairs, leveraging LLM agents to generate code and cross-validate semantic correctness. Experimental results demonstrate that this architecture substantially enhances scalability and trustworthiness—measured by joint coverage, branch consistency, certified unreachability, and escape rate—while decoupling trustworthiness from authorship.
📝 Abstract
Verified translation has two well-studied extremes: prove the translator once (certified compilation), or validate each run of one translator (translation validation). Both treat a single translation in isolation. We study translations as a graph -- many source languages, several reasoning targets, multiple independently built routes -- where the honest answer to "is this translation correct?" differs from edge to edge. We present a calculus of fidelity-graded translations: pairs of languages close commuting squares that are checkable per program and compose by pasting; declared fidelity grades compose by weakest link, are re-established per run by inline checking, and are exceeded by agreement between independently derived routes; and an end-to-end theorem isolates a fundamental asymmetry -- witness-carrying answers are self-certifying at the source, while universal answers are where grades, branches, and certificates earn their cost. The compositional core is mechanized in Lean 4. The calculus is implemented in hurdy-gurdy, a platform of 13 languages and 13 pairs around two reasoning hubs, built as a two-directional experiment in LLM-generated correctness: independent LLM agents wrote every pair, largely unsupervised, with the architecture's cross-checks as the only semantic gate, and the platform's intended player is itself an LLM. All code, and most of this paper, is LLM-generated; the human contribution is the architecture. The same gate is the intended growth model: hurdy-gurdy scales in language support through pairs contributed by anyone -- with LLMs, with agents, or by hand -- admitted by architecture, not authorship. We report conjoined coverage, branch agreement, a compliance-derived benchmark with machine-derived ground truth, witness replay, certified unreachability, and escape-rate experiments for the gate.
Problem

Research questions and friction points this paper is trying to address.

fidelity-graded translations
translation validation
untrusted authors
compositional verification
multi-language translation graph
Innovation

Methods, ideas, or system contributions that make the work stand out.

fidelity-graded translations
translation validation
compositional verification
LLM-generated correctness
cross-checking architecture
🔎 Similar Papers
No similar papers found.
C
Christoph Kirsch
University of Salzburg, Austria; Czech Technical University, Prague, Czechia