Institution profile

LightOn

Industry researcheurope · fr
Official website
Research library14linked papers
Opportunities0open roles
Selected work

Representative Papers

LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR

Jan 20, 2026

This work proposes a billion-parameter, end-to-end multilingual vision-language model that directly converts document images into well-structured, naturally ordered text while enabling precise localization of embedded images. Addressing the error-proneness and inefficiency of traditional multi-stage OCR pipelines in handling multilingual documents, the model introduces a novel curriculum-based bounding box localization strategy during pretraining. It further incorporates reinforcement learning with an IoU-based reward mechanism (RLVR) and enhances robustness through checkpoint averaging and task arithmetic fusion. Evaluated on OlmOCR-Bench, the model achieves state-of-the-art performance while being nine times smaller than the previous best model and offering significantly faster inference. The authors publicly release the model, training data, and a new evaluation benchmark, LightOnOCR-bbox-bench.

1 citationsRead paper

TrustMI: Causally controlling how assistants trust their users

Oct 05, 2026

This study addresses the security vulnerability arising from erroneous trust in large language model (LLM) assistants when they fail to verify user intent. To mitigate this, the authors construct contrastive dialogue datasets and identify linear directions within the activation space, enabling causal intervention on the model’s trust decisions through activation steering matrices while keeping parameters frozen. This work provides the first demonstration that LLM trust behavior can be controlled both monotonically and causally via such a mechanism. Furthermore, it achieves bidirectional trust regulation across multiple model families, effectively mitigating critical security threats, including harmful requests and prompt injection attacks.

0 citationsRead paper

Small Agents with Semantic Search: Efficient Multilingual Code Localization

Oct 04, 2026

This study addresses the inefficiency, high computational cost, and limited cross-lingual generalization of natural language file localization within code repositories. To overcome these challenges, we propose a lightweight agent-based approach leveraging ColGREP semantic retrieval. The method employs a late-interaction retrieval architecture and introduces a turn-level credit assignment training framework based on retrieval outcomes, achieving efficient file localization through weighted supervised fine-tuning combined with reinforcement learning. Experimental results demonstrate that the proposed approach reduces inference latency by 44.1% and token consumption by 29.1%, while significantly improving localization accuracy and generalization to unseen programming languages. These advancements effectively facilitate deployment on edge devices.

0 citationsRead paper

A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model

Sep 22, 2026

This study addresses the limitation of existing carbon footprint assessments for large language models, which predominantly focus on the training phase. Using the 13-billion-parameter Arabic model Noor as a case study, this work proposes a comprehensive life-cycle environmental impact assessment framework encompassing data processing, research and development, training, inference, and exogenous factors such as cross-border collaboration. By integrating zero-shot generalization, instruction fine-tuning, and distributed computing techniques, the authors conduct an end-to-end carbon emission accounting. The results demonstrate that inference and exogenous costs significantly influence the total carbon budget. This research transcends conventional assessment boundaries, providing empirical evidence and optimization pathways for mitigating the carbon footprint of extremely large-scale models.

0 citationsRead paper

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

Jul 29, 2026

This work addresses the poor reproducibility and limited multilingual transferability of existing retrieval models, which often rely on closed-source data. We propose a fully open-source, end-to-end training framework leveraging a reconstructed dataset of 665 million English contrastive pairs and 1.88 million supervised pairs to train both DenseOn (a single-vector dense retriever) and LateOn (a ColBERT-style late-interaction model). These models are extended to eight languages, marking the first public release of large-scale multilingual retrieval data and models. Experimental results show that LateOn significantly outperforms DenseOn on unseen languages, achieving average nDCG@10 scores of 56.20 and 57.22 respectively on BEIR—setting new state-of-the-art results at this scale. Furthermore, our analysis reveals that translate-train functions as a general multilingual generalization mechanism rather than merely a target-language expansion strategy. All code, data, and models are publicly released.

0 citationsRead paper
Recent publications

Latest Papers

TrustMI: Causally controlling how assistants trust their users

Oct 05, 2026

This study addresses the security vulnerability arising from erroneous trust in large language model (LLM) assistants when they fail to verify user intent. To mitigate this, the authors construct contrastive dialogue datasets and identify linear directions within the activation space, enabling causal intervention on the model’s trust decisions through activation steering matrices while keeping parameters frozen. This work provides the first demonstration that LLM trust behavior can be controlled both monotonically and causally via such a mechanism. Furthermore, it achieves bidirectional trust regulation across multiple model families, effectively mitigating critical security threats, including harmful requests and prompt injection attacks.

0 citationsRead paper

Small Agents with Semantic Search: Efficient Multilingual Code Localization

Oct 04, 2026

This study addresses the inefficiency, high computational cost, and limited cross-lingual generalization of natural language file localization within code repositories. To overcome these challenges, we propose a lightweight agent-based approach leveraging ColGREP semantic retrieval. The method employs a late-interaction retrieval architecture and introduces a turn-level credit assignment training framework based on retrieval outcomes, achieving efficient file localization through weighted supervised fine-tuning combined with reinforcement learning. Experimental results demonstrate that the proposed approach reduces inference latency by 44.1% and token consumption by 29.1%, while significantly improving localization accuracy and generalization to unseen programming languages. These advancements effectively facilitate deployment on edge devices.

0 citationsRead paper

A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model

Sep 22, 2026

This study addresses the limitation of existing carbon footprint assessments for large language models, which predominantly focus on the training phase. Using the 13-billion-parameter Arabic model Noor as a case study, this work proposes a comprehensive life-cycle environmental impact assessment framework encompassing data processing, research and development, training, inference, and exogenous factors such as cross-border collaboration. By integrating zero-shot generalization, instruction fine-tuning, and distributed computing techniques, the authors conduct an end-to-end carbon emission accounting. The results demonstrate that inference and exogenous costs significantly influence the total carbon budget. This research transcends conventional assessment boundaries, providing empirical evidence and optimization pathways for mitigating the carbon footprint of extremely large-scale models.

0 citationsRead paper

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

Jul 29, 2026

This work addresses the poor reproducibility and limited multilingual transferability of existing retrieval models, which often rely on closed-source data. We propose a fully open-source, end-to-end training framework leveraging a reconstructed dataset of 665 million English contrastive pairs and 1.88 million supervised pairs to train both DenseOn (a single-vector dense retriever) and LateOn (a ColBERT-style late-interaction model). These models are extended to eight languages, marking the first public release of large-scale multilingual retrieval data and models. Experimental results show that LateOn significantly outperforms DenseOn on unseen languages, achieving average nDCG@10 scores of 56.20 and 57.22 respectively on BEIR—setting new state-of-the-art results at this scale. Furthermore, our analysis reveals that translate-train functions as a general multilingual generalization mechanism rather than merely a target-language expansion strategy. All code, data, and models are publicly released.

0 citationsRead paper

Rethinking the Multilingual Reasoning Gap with Layer Swap

May 26, 2026

This study addresses the significant performance degradation of multilingual large language models in non-English reasoning, where maintaining both target-language chain-of-thought and reasoning accuracy remains challenging. The authors construct a six-language long-reasoning dataset and train native-language reasoning and English-pivot models based on Qwen3-8B-Base. Through weight-space analysis, they identify that core reasoning capabilities are concentrated in intermediate layers. Leveraging this insight, they propose an innovative Layer Swap method that exchanges intermediate-layer parameters to enhance native-language reasoning. Experiments demonstrate that this approach reduces the average reasoning gap to 1.9–3.5% across five non-English languages, effectively closing the performance gap while preserving target-language chain-of-thought throughout. This work further reveals, for the first time, a language-agnostic reasoning core alongside language-specific peripheral layer structures.

0 citationsRead paper