Institution profile

Prime Intellect

Industry research
Official website
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

Oct 08, 2026

In the evaluation of large language model (LLM) agents, fluctuations in verifier scores frequently conflate genuine capability changes with assessment bias. This work proposes the TRACE protocol, which systematically modifies individual evaluation components and conducts paired comparative runs to transform score variations into a testable causal diagnostic process, thereby precisely disentangling agent behavioral shifts from scoring rule artifacts. Through experiments involving synthetic tasks, public benchmarks, and repeated multi-agent trials, this study reveals that tool renaming induces spurious score degradations and high variance. The results demonstrate that TRACE effectively identifies measurement errors, offering a reliable attribution analysis framework for robust agent evaluation.

0 citationsRead paper

MirrorCode: AI can rebuild entire programs from behavior alone

Jun 29, 2026

Current AI programming evaluations are largely confined to short, isolated tasks and lack standardized benchmarks for assessing the ability to reproduce complete software systems end-to-end. This work proposes MirrorCode, a novel long-horizon programming benchmark that introduces a behavior-based reverse-engineering paradigm: under strict black-box conditions without access to source code, AI systems must precisely reconstruct the functionality of real-world software solely from its observable behavior. The benchmark encompasses 25 complex projects spanning Unix utilities, compilers, and bioinformatics tools. Leveraging large language model–driven autonomous agents and a high-fidelity test-validation framework, the strongest model achieves a 56% success rate, including the successful reproduction of large-scale tools such as gotree (16,000 lines), demonstrating a significant leap in AI’s capacity for complex software engineering tasks.

0 citationsRead paper

Arcee Trinity Large Technical Report

Feb 18, 2026

This work proposes the Trinity series of sparse Mixture-of-Experts (MoE) models—Trinity-Large, -Mini, and -Nano—to enhance parameter efficiency and training stability in large language models. The architecture integrates interleaved local-global attention, gated attention, depth-scaled Sandwich normalization, and a Sigmoid-based routing mechanism, and is trained using the Muon optimizer. A key innovation is the introduction of Soft-clamped Momentum Expert Bias Updates (SMEBU), a novel load-balancing strategy that significantly improves MoE training stability. All variants complete training without any loss spikes: Trinity-Nano and -Mini are pretrained on 10 trillion tokens, while Trinity-Large uses 17 trillion tokens. The codebase has been publicly released.

0 citationsRead paper
Recent publications

Latest Papers

TRACE: Diagnosing Verifier Brittleness in Agentic Evaluation

Oct 08, 2026

In the evaluation of large language model (LLM) agents, fluctuations in verifier scores frequently conflate genuine capability changes with assessment bias. This work proposes the TRACE protocol, which systematically modifies individual evaluation components and conducts paired comparative runs to transform score variations into a testable causal diagnostic process, thereby precisely disentangling agent behavioral shifts from scoring rule artifacts. Through experiments involving synthetic tasks, public benchmarks, and repeated multi-agent trials, this study reveals that tool renaming induces spurious score degradations and high variance. The results demonstrate that TRACE effectively identifies measurement errors, offering a reliable attribution analysis framework for robust agent evaluation.

0 citationsRead paper

MirrorCode: AI can rebuild entire programs from behavior alone

Jun 29, 2026

Current AI programming evaluations are largely confined to short, isolated tasks and lack standardized benchmarks for assessing the ability to reproduce complete software systems end-to-end. This work proposes MirrorCode, a novel long-horizon programming benchmark that introduces a behavior-based reverse-engineering paradigm: under strict black-box conditions without access to source code, AI systems must precisely reconstruct the functionality of real-world software solely from its observable behavior. The benchmark encompasses 25 complex projects spanning Unix utilities, compilers, and bioinformatics tools. Leveraging large language model–driven autonomous agents and a high-fidelity test-validation framework, the strongest model achieves a 56% success rate, including the successful reproduction of large-scale tools such as gotree (16,000 lines), demonstrating a significant leap in AI’s capacity for complex software engineering tasks.

0 citationsRead paper

Arcee Trinity Large Technical Report

Feb 18, 2026

This work proposes the Trinity series of sparse Mixture-of-Experts (MoE) models—Trinity-Large, -Mini, and -Nano—to enhance parameter efficiency and training stability in large language models. The architecture integrates interleaved local-global attention, gated attention, depth-scaled Sandwich normalization, and a Sigmoid-based routing mechanism, and is trained using the Muon optimizer. A key innovation is the introduction of Soft-clamped Momentum Expert Bias Updates (SMEBU), a novel load-balancing strategy that significantly improves MoE training stability. All variants complete training without any loss spikes: Trinity-Nano and -Mini are pretrained on 10 trillion tokens, while Trinity-Large uses 17 trillion tokens. The codebase has been publicly released.

0 citationsRead paper