🤖 AI Summary
This study addresses the absence of unified, open artificial intelligence resources in the telecommunications domain by proposing OTel, an open-source framework. By integrating multi-source data and evaluation partitions, this work constructs high-quality datasets encompassing tasks such as retrieval and reranking. Furthermore, it establishes a reproducible baseline for telecommunications AI development through full-parameter post-training, embedding models, and context-grounded large language models (LLMs). The project releases thirty post-trained models to facilitate community-driven extension. Experimental results demonstrate strong performance, achieving an NDCG@10 of 93.1% for retrieval, an MRR@10 of 0.947 for reranking, and an LLM accuracy of 87.8%. With cumulative downloads exceeding sixteen million, OTel provides a standardized and open foundation for advancing telecommunications AI research.
📝 Abstract
We present Open Telco (OTel), an open telecom AI resource that releases derived telecom datasets for retrieval, reranking, instruction tuning, and safety/abstention, together with 30 full-parameter post-trained baselines spanning 10 embedding models, 3 rerankers, and 17 language models. The community has already engaged substantially with the resource: as of May 3, 2026, the released models have been downloaded over 16 million times and the project has received 157+ pieces of media coverage worldwide. Building on prior open telecom datasets and benchmarks, OTel provides documented telecom data sources, held-out evaluation partitions, trained embedding models, rerankers, context-grounded LLMs, and safety/abstention data in one unified resource. Each baseline starts from an open-weight model and is post-trained on OTel-derived data using an open training recipe, then evaluated on held-out OTel evaluation partitions. OTel post-training improves performance across all three model families: embedding retrieval reaches 93.1% NDCG@10, reranking reaches 0.947 MRR@10, and language-model correctness reaches 87.8%. We release OTel as a reproducible starting point and invite the community to expand the data, improve embedding and reranking models, and build stronger context-grounded telecom LLMs.