OTel: Open Telco AI Datasets, Benchmarks, and Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of unified, open artificial intelligence resources in the telecommunications domain by proposing OTel, an open-source framework. By integrating multi-source data and evaluation partitions, this work constructs high-quality datasets encompassing tasks such as retrieval and reranking. Furthermore, it establishes a reproducible baseline for telecommunications AI development through full-parameter post-training, embedding models, and context-grounded large language models (LLMs). The project releases thirty post-trained models to facilitate community-driven extension. Experimental results demonstrate strong performance, achieving an NDCG@10 of 93.1% for retrieval, an MRR@10 of 0.947 for reranking, and an LLM accuracy of 87.8%. With cumulative downloads exceeding sixteen million, OTel provides a standardized and open foundation for advancing telecommunications AI research.
📝 Abstract
We present Open Telco (OTel), an open telecom AI resource that releases derived telecom datasets for retrieval, reranking, instruction tuning, and safety/abstention, together with 30 full-parameter post-trained baselines spanning 10 embedding models, 3 rerankers, and 17 language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. Each baseline starts from an open-weight model and is post-trained on OTel-derived data using an open training recipe, then evaluated on held-out OTel evaluation partitions. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.1% NDCG@10, reranking reaches 0.947 MRR@10, and language-model correctness reaches 87.8%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.
Problem

Research questions and friction points this paper is trying to address.

Telecom AI
Open Datasets
Benchmarks
Baseline Models
Reproducibility
Innovation

Methods, ideas, or system contributions that make the work stand out.

Telecom AI
Open Datasets
Post-training
Embedding Models
Reranking
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.