From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of translating abstract normative frameworks into trainable alignment data for large language models. Specifically, it presents the first systematic effort to convert Islamic ethical principles into Arabic and English supervised fine-tuning (SFT) and preference datasets. Employing an expert-driven methodology, the proposed approach integrates SFT, direct preference optimization (DPO), and blind evaluation techniques to achieve model alignment. The results validate the feasibility of leveraging religious ethical norms to guide LLM alignment. Notably, the SFT-aligned model attained a 51.3% expert preference rate, significantly enhancing output normativity without compromising general-purpose capabilities. By demonstrating that culturally grounded ethical frameworks can effectively steer model behavior, this work establishes a novel paradigm for cross-cultural value alignment in artificial intelligence.
📝 Abstract
Aligning language models with a specified normative framework requires translating abstract principles into concrete examples and preference signals from which models can learn. We present an expert-driven methodology for constructing such alignment data and apply it to a normative framework grounded in Islamic ethical, theological, and jurisprudential traditions. Over approximately one year, seven domain experts systematically probed language models to identify alignment deficiencies, curated desired responses, and constructed preference pairs from model outputs and expert judgments. The resulting Arabic-English datasets contain approximately 2.8K supervised fine-tuning (SFT) examples and 5.4K preference pairs spanning a broad range of normative domains. We evaluate the datasets through controlled post-training experiments comparing a Baseline model with models incorporating the curated SFT data alone and both the SFT and preference data. In blind expert evaluation on 150 separately constructed prompts, the model trained with the curated SFT data was preferred over the Baseline in 51.3% of assessor judgments, compared with 14.4% in the opposite direction (p<.001 at the prompt level). Adding the preference data resulted in a smaller difference, with the model trained with both datasets preferred over the SFT model in 28.0% of judgments versus 20.9% in the opposite direction; this difference was not statistically significant at the prompt level (p = .166). Standard Arabic and English benchmarks show no broad degradation in general-purpose capabilities. These results demonstrate how expert-defined normative principles can be systematically operationalized into alignment data and evaluated through controlled model training.
Problem

Research questions and friction points this paper is trying to address.

Language Model Alignment
Normative Frameworks
Supervised Fine-Tuning
Preference Data
AI Ethics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Normative Alignment
Supervised Fine-Tuning (SFT)
Preference Data
Expert-driven Methodology
Controlled Post-training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Husrev Taha Sencar
Husrev Taha Sencar
Qatar Computing Research Institute, HBKU
ai safety and securitythreat intelligencesecuritydigital forensicsmultimedia security
R
Rezart Beka
College of Islamic Studies, HBKU, Qatar
D
Danish Naeem
Argumentation and Conflict Studies, Ibn Haldun University, Turkiye
S
Seda Ozalkan
College of Islamic Studies, HBKU, Qatar
Majd Hawasly
Majd Hawasly
QCRI, Hamad Bin Khalifa University
Autonomous systemsLifelong learningNatural Language Processing
J
Ji Lucas
Qatar Computing Research Institute, HBKU, Qatar
A
Ala AlFuqaha
College of Science and Engineering, HBKU, Qatar
Mohamed Abdallah
Mohamed Abdallah
Professor and Associate Dean, Hamad Bin Khalifa University
Wireless CommunicationsEdge AIAutonomous VehiclesWireless SecuritySmart Grids
R
Recep Senturk
College of Islamic Studies, HBKU, Qatar