Institution profile

Graphcore

Industry researcheurope · gb
Official website
Research library16linked papers
Opportunities0open roles
Selected work

Representative Papers

UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG

Jan 28, 2026arXiv.org

This study addresses the challenges of hallucination in large language models (LLMs) and multi-hop reasoning retrieval over knowledge graphs (KGs) by proposing a training-free, generalizable KG-RAG framework. The method achieves efficient knowledge graph question answering through the synergistic integration of LLM-based query generation, an inductive neural executor, and an arbitration mechanism. Its core innovation lies in introducing the first training-agnostic paradigm capable of scaling to Wikidata-level massive knowledge graphs without requiring model retraining. Experimental results demonstrate that the proposed approach attains state-of-the-art performance on KGQA benchmarks while significantly reducing inference costs and enhancing factual accuracy.

1 citationsRead paper

Studying quantization trade-offs for efficient inference deployment in machine translation

Jul 31, 2026

This study addresses the lack of systematic evaluation of how quantization impacts inference efficiency and translation quality under orchestrated workloads in real-world server deployments of large language models, particularly given that mainstream benchmarks overlook long-context dynamics. We present the first controlled assessment of EuroLLM and Hy-MT2 (1.7B–22B parameters) on A100/H100 GPUs, combining W4A8/W8A8 quantization with document chunking strategies. Introducing a document-level evaluation based on WMT24++, we uncover quantization–long-context interactions invisible to sentence-level benchmarks. Our experiments reveal that Hy-MT2 maintains robustness under quantization, whereas EuroLLM suffers significant quality degradation. Jointly optimizing chunking and quantization substantially improves the latency–throughput Pareto frontier, demonstrating that the trade-off between translation quality and efficiency is highly dependent on model architecture, quantization format, and chunking strategy.

0 citationsRead paper
Recent publications

Latest Papers

Studying quantization trade-offs for efficient inference deployment in machine translation

Jul 31, 2026

This study addresses the lack of systematic evaluation of how quantization impacts inference efficiency and translation quality under orchestrated workloads in real-world server deployments of large language models, particularly given that mainstream benchmarks overlook long-context dynamics. We present the first controlled assessment of EuroLLM and Hy-MT2 (1.7B–22B parameters) on A100/H100 GPUs, combining W4A8/W8A8 quantization with document chunking strategies. Introducing a document-level evaluation based on WMT24++, we uncover quantization–long-context interactions invisible to sentence-level benchmarks. Our experiments reveal that Hy-MT2 maintains robustness under quantization, whereas EuroLLM suffers significant quality degradation. Jointly optimizing chunking and quantization substantially improves the latency–throughput Pareto frontier, demonstrating that the trade-off between translation quality and efficiency is highly dependent on model architecture, quantization format, and chunking strategy.

0 citationsRead paper

A Practical Investigation of Training-free Relaxed Speculative Decoding

Jul 09, 2026

This work investigates how to accelerate large language model inference by relaxing the strict fidelity constraints of conventional speculative decoding without compromising generation quality. We present the first systematic evaluation of various training-free relaxed speculative decoding strategies, unifying existing frameworks and benchmarking them on modern large models. Our analysis reveals that most relaxation methods heavily rely on the draft model’s language modeling capabilities and struggle to generalize to lightweight, specialized predictors. Nevertheless, well-designed relaxation mechanisms can achieve a controllable trade-off between speed and capability—and may even yield modest performance gains. This study distills practical insights for practitioners and underscores the critical role of draft model capability assessment in effective relaxed decoding.

0 citationsRead paper