Institution profile

Emergence AI

Industry researchnorthamerica · us
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond Type-checking: Towards Holistic Evaluation of Formal Specification Generation

Oct 06, 2026

This study addresses the lack of deterministic verification of user intent in formal specification generation, which risks producing proofs grounded in erroneous specifications. To this end, it constructs a unified dataset and multidimensional evaluation framework encompassing formal validity, similarity, and behavioral adequacy. The work proposes a strategy distinguishing input acceptance from output constraints to clarify the evidential scope of metrics, and integrates an LLM agent workflow, the Lean theorem prover, and Generalized Tree Edit Distance (GTED) to enable automated evaluation. The findings reveal the limitations of single similarity metrics and the impact of metric coverage on ranking outcomes, demonstrating that perfect postcondition scores may obscure deficiencies in input contracts. Ultimately, this research provides a systematic benchmark for evaluating formal specifications.

0 citationsRead paper

Second-Order Problem Solving for Recursive Self-Improvement in Formal Verification

Oct 05, 2026

This study addresses the trial-and-error oscillations and residual latent defects caused by blind parameter patching in recursive self-improvement (RSI) by proposing the SO-RSI framework. This method employs passive execution trajectory monitoring to detect structural anomalies and trigger lightweight diagnostic probes, while accumulating causal evidence through cross-round persistent inquiry memory. Consequently, optimization is elevated to second-order diagnostic investigation that guides systematic workflow editing, achieving a paradigm shift from symptom response to mechanistic understanding. Formal verification using Lean 4 and Verus demonstrates that, under equivalent search budgets, pass rates for proof and code generation improve by 21.8 and 25.8 percentage points, respectively. These results confirm that SO-RSI effectively suppresses failure recurrence and eliminates futile optimization loops.

0 citationsRead paper

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy

Jun 06, 2026

This study addresses the limitations of current evaluations of large language model (LLM) agents, which are predominantly confined to short-term, static tasks and fail to capture dynamic phenomena such as behavioral drift, inter-agent interactions, and emergent governance over extended periods. To bridge this gap, we introduce a persistent simulation platform enabling long-term coexistence of heterogeneous LLM agents, integrating real-time external data (e.g., weather, news), over 120 specialized tools, three-tiered persistent memory, and a democratic decision-making mechanism. For the first time, this framework supports measurable analysis of multi-agent behavioral evolution, coordination patterns, and governance structures at weekly to monthly timescales. A 15-day cross-vendor experiment across five parallel worlds revealed stark divergences—from stable self-governance to collective collapse—and we release all prompts, logs, and configurations to foster reproducible research.

0 citationsRead paper

Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents

Sep 13, 2025

This work addresses the longstanding issue in automatic summarization—overemphasis on content coverage at the expense of reader engagement. To this end, we propose the “Spotlight” paradigm, which prioritizes extracting and generating the most salient, attention-grabbing information to enhance reader involvement with the source text. We formally define the “spotlight” concept—distinct from conventional summarization—and introduce the first dedicated dataset and evaluation benchmark explicitly designed for measuring reader engagement. Our method employs a two-stage training strategy: first, fine-tuning a large language model on our curated dataset; second, aligning outputs with user preferences via Direct Preference Optimization (DPO). Extensive experiments demonstrate that our approach significantly outperforms baseline summarization models across key dimensions—including salient information identification, readability, and perceptual appeal—establishing a novel, engagement-centered standard for information distillation.

0 citationsRead paper
Recent publications

Latest Papers

Beyond Type-checking: Towards Holistic Evaluation of Formal Specification Generation

Oct 06, 2026

This study addresses the lack of deterministic verification of user intent in formal specification generation, which risks producing proofs grounded in erroneous specifications. To this end, it constructs a unified dataset and multidimensional evaluation framework encompassing formal validity, similarity, and behavioral adequacy. The work proposes a strategy distinguishing input acceptance from output constraints to clarify the evidential scope of metrics, and integrates an LLM agent workflow, the Lean theorem prover, and Generalized Tree Edit Distance (GTED) to enable automated evaluation. The findings reveal the limitations of single similarity metrics and the impact of metric coverage on ranking outcomes, demonstrating that perfect postcondition scores may obscure deficiencies in input contracts. Ultimately, this research provides a systematic benchmark for evaluating formal specifications.

0 citationsRead paper

Second-Order Problem Solving for Recursive Self-Improvement in Formal Verification

Oct 05, 2026

This study addresses the trial-and-error oscillations and residual latent defects caused by blind parameter patching in recursive self-improvement (RSI) by proposing the SO-RSI framework. This method employs passive execution trajectory monitoring to detect structural anomalies and trigger lightweight diagnostic probes, while accumulating causal evidence through cross-round persistent inquiry memory. Consequently, optimization is elevated to second-order diagnostic investigation that guides systematic workflow editing, achieving a paradigm shift from symptom response to mechanistic understanding. Formal verification using Lean 4 and Verus demonstrates that, under equivalent search budgets, pass rates for proof and code generation improve by 21.8 and 25.8 percentage points, respectively. These results confirm that SO-RSI effectively suppresses failure recurrence and eliminates futile optimization loops.

0 citationsRead paper

Emergence World: A Platform for Evaluating Long-Horizon Multi-Agent Autonomy

Jun 06, 2026

This study addresses the limitations of current evaluations of large language model (LLM) agents, which are predominantly confined to short-term, static tasks and fail to capture dynamic phenomena such as behavioral drift, inter-agent interactions, and emergent governance over extended periods. To bridge this gap, we introduce a persistent simulation platform enabling long-term coexistence of heterogeneous LLM agents, integrating real-time external data (e.g., weather, news), over 120 specialized tools, three-tiered persistent memory, and a democratic decision-making mechanism. For the first time, this framework supports measurable analysis of multi-agent behavioral evolution, coordination patterns, and governance structures at weekly to monthly timescales. A 15-day cross-vendor experiment across five parallel worlds revealed stark divergences—from stable self-governance to collective collapse—and we release all prompts, logs, and configurations to foster reproducible research.

0 citationsRead paper

Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents

Sep 13, 2025

This work addresses the longstanding issue in automatic summarization—overemphasis on content coverage at the expense of reader engagement. To this end, we propose the “Spotlight” paradigm, which prioritizes extracting and generating the most salient, attention-grabbing information to enhance reader involvement with the source text. We formally define the “spotlight” concept—distinct from conventional summarization—and introduce the first dedicated dataset and evaluation benchmark explicitly designed for measuring reader engagement. Our method employs a two-stage training strategy: first, fine-tuning a large language model on our curated dataset; second, aligning outputs with user preferences via Direct Preference Optimization (DPO). Extensive experiments demonstrate that our approach significantly outperforms baseline summarization models across key dimensions—including salient information identification, readability, and perceptual appeal—establishing a novel, engagement-centered standard for information distillation.

0 citationsRead paper