Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization

📅 2026-07-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Existing methods struggle to efficiently and accurately evaluate the similarity between structured JSON outputs generated by large language models and target schemas, particularly when handling graph structures and identifier renaming. This work proposes Object Aligner—a deterministic, schema-driven similarity scoring tool that recursively aligns tree structures, supports graph reference relationships, and integrates the Hungarian algorithm for unordered sets, sequence alignment for ordered structures, and an approximation of Weisfeiler–Leman graph coloring to compute a bijection over identifiers. To our knowledge, this is the first approach enabling fine-grained, interpretable, and renaming-invariant matching for graph-structured JSON with arbitrary identifiers. Deployed as a reward signal in the GEPA prompt optimizer, it consistently matches or improves performance across multiple datasets while providing precise error localization and repair suggestions at zero additional computational cost.
📝 Abstract
Large language models (LLMs) are often asked to produce JSON conforming to a fixed schema, powering information extraction, tool calling, agentic planning, and knowledge-graph construction. Measuring how closely an output matches a gold reference is essential yet surprisingly hard: exact match is brittle, text similarity ignores structure, and an LLM judge is expensive, opaque, and non-deterministic. We address this with Object Aligner (OA), an open-source Python library that scores two JSON objects deterministically by recursively aligning their trees (the Hungarian algorithm for unordered collections, sequence alignment for ordered ones) and awarding partial credit at the granularity the schema declares. The Object Aligner is configured entirely through a set of JSON Schema extensions, so adapting it to a new task involves annotating a schema rather than writing code. Complex structured data, however, are rarely flat trees: records may form graphs or hypergraphs keyed by arbitrary identifiers, breaking the assumptions of prior similarity metrics. Our central contribution, referential alignment, closes this gap by inferring a bijection between gold and candidate identifiers and scoring every reference through it, so the score is invariant to relabeling. Since recovering this bijection exactly is graph isomorphism, the Object Aligner approximates it with Weisfeiler-Leman color refinement. An order-sensitive sequence regime targets ranking and planning. Since the same alignment localizes every mismatch, the Object Aligner emits ranked repair suggestions at no extra cost. Used as a reward inside the GEPA prompt optimizer, Object Aligner helps or stays neutral across all datasets.
Problem

Research questions and friction points this paper is trying to address.

JSON schema similarity
structured output evaluation
graph alignment
referential integrity
LLM output scoring
Innovation

Methods, ideas, or system contributions that make the work stand out.

Object Aligner
JSON Schema
referential alignment
graph isomorphism
prompt optimization
🔎 Similar Papers
No similar papers found.