Institution profile

Ekimetrics

Industry researcheurope · fr
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

TICDA: Tabular In-Context Data Attribution

Oct 06, 2026

This study addresses the challenge of efficiently quantifying the influence of demonstration examples on predictions during in-context learning with tabular foundation models, where conventional attribution methods encounter significant computational bottlenecks. To this end, this work proposes TICDA, a framework that directly measures the data attribution contribution of demonstrations to prediction outcomes by constructing linear surrogate models and conducting latent space embedding analysis. Notably, this approach requires only a single forward pass, eliminating the need for parameter updates or multiple inference iterations. The proposed TICDA framework achieves low-cost, high-precision data influence quantification while effectively balancing attribution accuracy with inference efficiency. Extensive evaluations demonstrate its superior performance across diverse downstream tasks, including error detection, context selection, and active learning, establishing it as a practical solution for interpretable in-context learning in tabular domains.

0 citationsRead paper

TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

May 07, 2026

This work addresses the inflexibility of existing tabular foundation models in adapting to downstream tasks during inference, as conventional fine-tuning or parameter-efficient methods incur substantial computational overhead and rely heavily on internal model architecture. To overcome these limitations, we propose a lightweight, architecture-agnostic input-space residual adapter that operates under a frozen backbone. The adapter learns task-specific input perturbations through end-to-end training and incorporates an identity fallback mechanism, allowing the validation set to automatically determine whether adaptation should be activated—thus balancing performance and robustness. Without modifying any model weights, our approach achieves significant gains on TabArena-Lite, with TabICLv2-Retouche surpassing the baseline by +56 Elo points and attaining a Pareto-optimal trade-off between predictive quality and training/inference efficiency.

0 citationsRead paper

Adaptive Chunking: Optimizing Chunking-Method Selection for RAG

Mar 26, 2026

This work addresses the limitations of traditional RAG systems, which rely on fixed chunking strategies ill-suited for diverse document structures and lack task-agnostic metrics to evaluate chunk quality. The authors propose the first adaptive chunking framework that dynamically selects the optimal chunking method based on document characteristics. They introduce five novel document-level intrinsic metrics—such as References Completeness and Intrachunk Cohesion—to guide chunking strategy selection without requiring downstream task feedback. The framework integrates an LLM-regex chunker and a recursive merging chunker, augmented with post-processing techniques. Evaluated across multiple domains without any model or prompt tuning, the approach improves QA accuracy from 62–64% to 72% and increases the number of correctly answered questions by over 30% (from 49 to 65).

0 citationsRead paper

Oh That Looks Familiar: A Novel Similarity Measure for Spreadsheet Template Discovery

Nov 10, 2025

Traditional spreadsheet template identification suffers from poor distinguishability due to high similarity in both spatial layout and data-type patterns. To address this, we propose a fine-grained similarity metric that jointly encodes semantic embeddings, data-type representations, and cell-level spatial coordinates. Our method is the first to integrate Chamfer and Hausdorff distances in an unsupervised framework, enabling holistic modeling of semantic, typological, and geometric information. Operating at the cell level, it achieves a perfect Adjusted Rand Index of 1.00 on the FUSTE benchmark—significantly outperforming the graph-based baseline Mondrian (0.90)—and enables exact template clustering and reconstruction. The approach supports downstream applications including retrieval-augmented generation and large-scale data cleaning. By delivering scalable, high-precision template discovery, it establishes a new paradigm for structured spreadsheet analysis.

0 citationsRead paper

Agentic RAG with Knowledge Graphs for Complex Multi-Hop Reasoning in Real-World Applications

Jul 22, 2025

Traditional RAG systems struggle with complex, multi-hop queries in knowledge-intensive domains—such as cross-entity association or author-wide document retrieval—due to their limited capacity for structured reasoning and semantic aggregation. To address this, we propose INRAExplorer, an agent-based RAG framework grounded in domain-specific knowledge graphs for agricultural, food, and environmental science literature. It integrates LLM agents with dynamic, multi-tool orchestration—including iterative retrieval, author-wide collection, and relational inference—as well as automated knowledge graph construction and multi-step reasoning algorithms. Our key contribution lies in the tight coupling of agent architecture with a curated, structured knowledge graph, enabling interpretable, graph-aware multi-hop question answering. Evaluated on real-world scientific corpora, INRAExplorer significantly improves answer completeness and accuracy for complex queries, while supporting high-level semantic search and cross-document information synthesis.

0 citationsRead paper
Recent publications

Latest Papers

TICDA: Tabular In-Context Data Attribution

Oct 06, 2026

This study addresses the challenge of efficiently quantifying the influence of demonstration examples on predictions during in-context learning with tabular foundation models, where conventional attribution methods encounter significant computational bottlenecks. To this end, this work proposes TICDA, a framework that directly measures the data attribution contribution of demonstrations to prediction outcomes by constructing linear surrogate models and conducting latent space embedding analysis. Notably, this approach requires only a single forward pass, eliminating the need for parameter updates or multiple inference iterations. The proposed TICDA framework achieves low-cost, high-precision data influence quantification while effectively balancing attribution accuracy with inference efficiency. Extensive evaluations demonstrate its superior performance across diverse downstream tasks, including error detection, context selection, and active learning, establishing it as a practical solution for interpretable in-context learning in tabular domains.

0 citationsRead paper

TFM-Retouche: A Lightweight Input-Space Adapter for Tabular Foundation Models

May 07, 2026

This work addresses the inflexibility of existing tabular foundation models in adapting to downstream tasks during inference, as conventional fine-tuning or parameter-efficient methods incur substantial computational overhead and rely heavily on internal model architecture. To overcome these limitations, we propose a lightweight, architecture-agnostic input-space residual adapter that operates under a frozen backbone. The adapter learns task-specific input perturbations through end-to-end training and incorporates an identity fallback mechanism, allowing the validation set to automatically determine whether adaptation should be activated—thus balancing performance and robustness. Without modifying any model weights, our approach achieves significant gains on TabArena-Lite, with TabICLv2-Retouche surpassing the baseline by +56 Elo points and attaining a Pareto-optimal trade-off between predictive quality and training/inference efficiency.

0 citationsRead paper

Adaptive Chunking: Optimizing Chunking-Method Selection for RAG

Mar 26, 2026

This work addresses the limitations of traditional RAG systems, which rely on fixed chunking strategies ill-suited for diverse document structures and lack task-agnostic metrics to evaluate chunk quality. The authors propose the first adaptive chunking framework that dynamically selects the optimal chunking method based on document characteristics. They introduce five novel document-level intrinsic metrics—such as References Completeness and Intrachunk Cohesion—to guide chunking strategy selection without requiring downstream task feedback. The framework integrates an LLM-regex chunker and a recursive merging chunker, augmented with post-processing techniques. Evaluated across multiple domains without any model or prompt tuning, the approach improves QA accuracy from 62–64% to 72% and increases the number of correctly answered questions by over 30% (from 49 to 65).

0 citationsRead paper

Oh That Looks Familiar: A Novel Similarity Measure for Spreadsheet Template Discovery

Nov 10, 2025

Traditional spreadsheet template identification suffers from poor distinguishability due to high similarity in both spatial layout and data-type patterns. To address this, we propose a fine-grained similarity metric that jointly encodes semantic embeddings, data-type representations, and cell-level spatial coordinates. Our method is the first to integrate Chamfer and Hausdorff distances in an unsupervised framework, enabling holistic modeling of semantic, typological, and geometric information. Operating at the cell level, it achieves a perfect Adjusted Rand Index of 1.00 on the FUSTE benchmark—significantly outperforming the graph-based baseline Mondrian (0.90)—and enables exact template clustering and reconstruction. The approach supports downstream applications including retrieval-augmented generation and large-scale data cleaning. By delivering scalable, high-precision template discovery, it establishes a new paradigm for structured spreadsheet analysis.

0 citationsRead paper

Agentic RAG with Knowledge Graphs for Complex Multi-Hop Reasoning in Real-World Applications

Jul 22, 2025

Traditional RAG systems struggle with complex, multi-hop queries in knowledge-intensive domains—such as cross-entity association or author-wide document retrieval—due to their limited capacity for structured reasoning and semantic aggregation. To address this, we propose INRAExplorer, an agent-based RAG framework grounded in domain-specific knowledge graphs for agricultural, food, and environmental science literature. It integrates LLM agents with dynamic, multi-tool orchestration—including iterative retrieval, author-wide collection, and relational inference—as well as automated knowledge graph construction and multi-step reasoning algorithms. Our key contribution lies in the tight coupling of agent architecture with a curated, structured knowledge graph, enabling interpretable, graph-aware multi-hop question answering. Evaluated on real-world scientific corpora, INRAExplorer significantly improves answer completeness and accuracy for complex queries, while supporting high-level semantic search and cross-document information synthesis.

0 citationsRead paper