Institution profile

MLCommons

Academic institutionnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

OTel: Open Telco AI Datasets, Benchmarks, and Models

Oct 06, 2026

This study addresses the absence of unified, open artificial intelligence resources in the telecommunications domain by proposing OTel, an open-source framework. By integrating multi-source data and evaluation partitions, this work constructs high-quality datasets encompassing tasks such as retrieval and reranking. Furthermore, it establishes a reproducible baseline for telecommunications AI development through full-parameter post-training, embedding models, and context-grounded large language models (LLMs). The project releases thirty post-trained models to facilitate community-driven extension. Experimental results demonstrate strong performance, achieving an NDCG@10 of 93.1% for retrieval, an MRR@10 of 0.947 for reranking, and an LLM accuracy of 87.8%. With cumulative downloads exceeding sixteen million, OTel provides a standardized and open foundation for advancing telecommunications AI research.

0 citationsRead paper

MLCommons Jailbreak Benchmark v1.0

Oct 02, 2026

This study addresses the lack of systematic quantitative evaluation methods for assessing the safety robustness of large language models under jailbreak attacks. To this end, we construct an end-to-end evaluation pipeline that employs a taxonomy-driven attack selection strategy based on the AILuminate benchmark, integrating human annotation with automated calibration mechanisms. Furthermore, we introduce a "resilience gap" metric to precisely quantify the divergence in safety performance between baseline and adversarial conditions. This work establishes a reproducible comparative benchmark alongside a risk disclosure framework. Experimental results demonstrate that models exhibit an average resilience gap of 7.57%, with unsafe response rates increasing significantly from 11.08% to 18.65%, thereby revealing substantial safety vulnerabilities in current large language models.

0 citationsRead paper
Recent publications

Latest Papers

OTel: Open Telco AI Datasets, Benchmarks, and Models

Oct 06, 2026

This study addresses the absence of unified, open artificial intelligence resources in the telecommunications domain by proposing OTel, an open-source framework. By integrating multi-source data and evaluation partitions, this work constructs high-quality datasets encompassing tasks such as retrieval and reranking. Furthermore, it establishes a reproducible baseline for telecommunications AI development through full-parameter post-training, embedding models, and context-grounded large language models (LLMs). The project releases thirty post-trained models to facilitate community-driven extension. Experimental results demonstrate strong performance, achieving an NDCG@10 of 93.1% for retrieval, an MRR@10 of 0.947 for reranking, and an LLM accuracy of 87.8%. With cumulative downloads exceeding sixteen million, OTel provides a standardized and open foundation for advancing telecommunications AI research.

0 citationsRead paper

MLCommons Jailbreak Benchmark v1.0

Oct 02, 2026

This study addresses the lack of systematic quantitative evaluation methods for assessing the safety robustness of large language models under jailbreak attacks. To this end, we construct an end-to-end evaluation pipeline that employs a taxonomy-driven attack selection strategy based on the AILuminate benchmark, integrating human annotation with automated calibration mechanisms. Furthermore, we introduce a "resilience gap" metric to precisely quantify the divergence in safety performance between baseline and adversarial conditions. This work establishes a reproducible comparative benchmark alongside a risk disclosure framework. Experimental results demonstrate that models exhibit an average resilience gap of 7.57%, with unsafe response rates increasing significantly from 11.08% to 18.65%, thereby revealing substantial safety vulnerabilities in current large language models.

0 citationsRead paper