Institution profile

Sirjan University of Technology

Academic institutionasia · ir
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning

Jul 22, 2026

This work proposes the Sentence Splitter framework, which formalizes the discovery of implicit factual structures in natural language sentences—specifically, descriptive prefixes and their corresponding factual completions—as a self-supervised learning task. Built upon the T5 architecture, the approach models sentence splitting as a discrete segmentation problem, generating self-supervised signals through templated symbolic pairs and recovering factual completions via probabilistic sequence generation. A lightweight bootstrapping mechanism further expands plausible prefix-completion structures. Without requiring human annotations, the method extracts structured prefix-tail pairs directly from raw text, effectively bridging symbolic knowledge with natural language. Experimental results demonstrate that the extracted supervision signals substantially enhance performance on downstream tasks such as knowledge graph completion and commonsense question answering, confirming the framework’s effectiveness and generalization capability.

0 citationsRead paper

The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

Jul 22, 2026

This work addresses the critical yet underexplored challenge of aligning prompt templates with the pretraining objectives of language models for knowledge generation tasks, where performance heavily depends on such alignment yet lacks effective guidance for template selection. To bridge this gap, the paper introduces the Maskedness Index (MI)—a novel metric that quantifies the degree of alignment between a given task and the model’s pretraining objective. Built upon the DepthRank algorithm, MI evaluates template suitability by measuring the discrepancy in knowledge relational scores between masked and prefix-style prompting formats. The proposed approach offers both theoretical grounding and a practical tool for prompt engineering, demonstrating on the ATOMIC2020 benchmark a significant positive correlation between MI and downstream generation performance, particularly yielding substantial gains in low-resource knowledge extraction scenarios.

0 citationsRead paper
Recent publications

Latest Papers

Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning

Jul 22, 2026

This work proposes the Sentence Splitter framework, which formalizes the discovery of implicit factual structures in natural language sentences—specifically, descriptive prefixes and their corresponding factual completions—as a self-supervised learning task. Built upon the T5 architecture, the approach models sentence splitting as a discrete segmentation problem, generating self-supervised signals through templated symbolic pairs and recovering factual completions via probabilistic sequence generation. A lightweight bootstrapping mechanism further expands plausible prefix-completion structures. Without requiring human annotations, the method extracts structured prefix-tail pairs directly from raw text, effectively bridging symbolic knowledge with natural language. Experimental results demonstrate that the extracted supervision signals substantially enhance performance on downstream tasks such as knowledge graph completion and commonsense question answering, confirming the framework’s effectiveness and generalization capability.

0 citationsRead paper

The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models

Jul 22, 2026

This work addresses the critical yet underexplored challenge of aligning prompt templates with the pretraining objectives of language models for knowledge generation tasks, where performance heavily depends on such alignment yet lacks effective guidance for template selection. To bridge this gap, the paper introduces the Maskedness Index (MI)—a novel metric that quantifies the degree of alignment between a given task and the model’s pretraining objective. Built upon the DepthRank algorithm, MI evaluates template suitability by measuring the discrepancy in knowledge relational scores between masked and prefix-style prompting formats. The proposed approach offers both theoretical grounding and a practical tool for prompt engineering, demonstrating on the ATOMIC2020 benchmark a significant positive correlation between MI and downstream generation performance, particularly yielding substantial gains in low-resource knowledge extraction scenarios.

0 citationsRead paper