feature extraction

Design, build, and evaluate feature-extraction pipelines and algorithms that compute and select quantitative descriptors from input data — including morphometric shape/size descriptors and stylometric linguistic/style metrics — to represent objects, documents, or signals for downstream modeling. Analyze and validate those features by quantifying corpus- or population-level differences, selecting predictive subsets, and assessing plausibility and limitations (e.g., given resolution or measurement constraints).

featureextraction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$191K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Palmistry-Informed Feature Extraction and Analysis using Machine Learning

Sep 02, 2025
SS
Shweta Samadhan Patil
🏛️ D.Y. Patil University

This study addresses the longstanding limitation in palmistry—its reliance on subjective interpretation and lack of quantitatively validated correlations—by proposing the first data-driven machine learning framework. Methodologically, it establishes an end-to-end computer vision pipeline integrating principal line extraction, texture modeling, and geometric morphology quantification, trained on a curated, expert-annotated dataset using supervised learning and optimized for lightweight mobile deployment. Its key contribution lies in the first systematic translation of culturally embedded palm features into reproducible, empirically verifiable numerical biomarkers, enabling statistically robust associations between palm morphology and externally observable phenotypic traits (e.g., physiological or behavioral characteristics). Experimental results demonstrate the model’s efficacy in discerning complex palm patterns, exhibiting high robustness and promising applicability in digital anthropometry and personalized phenotypic analysis.

Automated analysis of palmar features using machine learningExtracting key palm characteristics like lines and textureStudying correlations between palm morphology and traits

This study addresses the challenge that existing handwritten text recognition methods lack interpretable visual metrics suitable for paleographic analysis. The authors propose a novel architecture requiring only line-level transcription supervision, integrating a Transformer-based detection model, prototype-driven character representation learning, and a line-level reconstruction module to enable weakly supervised character localization and deformation modeling. This approach is the first to support automated paleographic measurements of individual characters, bigrams, and inter-character spacing. Evaluated on 160 pages from the 14th-century manuscript BnF fr. 2813, the method effectively distinguishes scribal styles and reveals subtle writing variations using only single-column text, significantly outperforming baselines such as Learnable Typewriter. Code and data are publicly released.

handwritten text recognitionhistorical scriptmorphological analysis

This work proposes a stylometric approach inspired by genome-wide association studies (GWAS), treating words as “genetic variants” and authorship as the “phenotype.” By applying logistic regression combined with multiple testing correction, the method identifies author-specific lexical markers that exhibit statistical significance in textual data. It represents the first adaptation of the GWAS paradigm to stylometric analysis, offering both statistical rigor and interpretability of results. Experimental validation across multilingual corpora—including English, German, and Russian—demonstrates the method’s ability to reliably detect stable, author-unique lexical features, thereby confirming its effectiveness and cross-lingual generalizability.

authorshipGWASinterpretability

Collage: Decomposable Rapid Prototyping for Information Extraction on Scientific PDFs

Oct 30, 2024
SG
Sireesh Gururaja
🏛️ Carnegie Mellon University

Scientific PDF information extraction tools suffer from inconsistent input formats, opaque “black-box” behavior, poor fault tolerance, and limited format support—hindering literature analysis efficiency for non-NLP researchers. To address these challenges, we propose the first modular information extraction experimental framework specifically designed for scientific PDFs, enabling model-level decoupling, fine-grained intermediate-state visualization, and unified cross-model evaluation. The framework integrates Hugging Face token classifiers, diverse large language models (LLMs), and domain-specific models within a PDF processing pipeline comprising layout-aware parsing, text reconstruction, and semantic alignment. Evaluated on materials science literature review tasks, it significantly reduces model trial-and-error overhead, improves error attribution accuracy, and enhances system interpretability. The platform delivers a debuggable, reusable, plug-and-play rapid prototyping capability for scientific information extraction, empowering domain experts without NLP expertise.

Challenges in debugging and understanding NLP processing pipelinesDifficulty comparing multimodal NLP models for scientific PDFsLack of tools for prototyping and evaluating extraction models

This work proposes a unified natural language processing framework to address key challenges in academic integrity, including plagiarism, content fabrication, and authorship verification. The framework integrates four core stylometric tasks: classification of human- versus machine-generated text, distinction between single- and multi-author documents, detection of authorship changes within multi-author texts, and identification of contributing authors in collaborative writing. The study introduces and publicly releases the first academic text dataset generated using Gemini under two distinct instruction settings—standard and strict—and systematically evaluates how prompting strategies affect detection performance. Experimental results demonstrate that texts produced under strict instructions are significantly more adversarial, thereby increasing the difficulty of accurate identification. The code and dataset are made openly available, establishing a new benchmark for research on academic integrity.

academic integrityauthorship verificationmachine-generated text

Latest Papers

What's happening recently
View more

This work addresses the challenge that domain experts face in translating natural language descriptions of data quality requirements into executable analyses, a process often hindered by reliance on data engineers, resulting in inefficiency and high technical barriers. To overcome this, the paper proposes a no-code, model-driven pipeline that leverages a QPM metamodel to define domain-specific quality analysis templates. Coupled with the Constrainify toolchain, it automatically transforms natural language requirements into executable and reusable analytical logic. By integrating model-driven engineering, metamodeling, and no-code web technologies, the approach significantly reduces dependency on technical expertise, enabling efficient, reproducible, and semantically aligned data quality assessments. This advancement enhances both the accessibility and automation of data quality analysis for non-technical domain practitioners.

data qualitydomain expertsno-code

This study addresses the longstanding reliance on subjective judgment in narrative quality assessment by introducing a computational framework grounded in 33 quantifiable linguistic features spanning lexical, syntactic, and semantic dimensions. For the first time, this work systematically applies multidimensional quantitative stylometric indicators to the automatic evaluation of narrative quality. Leveraging natural language processing, clustering analysis, and similarity matrix construction, the proposed model achieves near-perfect discrimination between texts authored by professional editors and self-published writers. Furthermore, it significantly outperforms existing evaluation metrics on a manually annotated dataset, thereby overcoming the limitations inherent in traditional story-level assessment approaches.

automatic assessmentlinguistic featuresnarrative evaluation

This study addresses the lack of efficient, automated methods for structuring heterogeneous real estate questionnaire documents. To this end, the authors propose an end-to-end information extraction framework that first categorizes documents into structural types using K-Means clustering and text classification. Subsequently, it leverages the DeepSeek-R1 large language model enhanced with prompt engineering to accurately extract 35 predefined attributes from complex document formats—including checkboxes and scanned images—and outputs them as structured JSON. Evaluated on a dataset of 2,781 documents, the method produced 2,766 unique property records. Downstream validation demonstrated a Jaccard similarity of 0.82, marking the first high-precision, scalable solution for structured information extraction from such challenging real estate documentation.

automated processingheterogeneous documentsproperty metadata

Hot Scholars

TW

Tianyang Wang

University of Alabama at Birmingham
machine learning (deep learning)computer vision
GC

Gal Chechik

NVIDIA, Bar Ilan University
Machine learningAIMachine perception
TC

Tianlong Chen

Assistant Professor, CS@UNC Chapel Hill; Chief AI Scientist, hireEZ
Machine LearningAI4ScienceComputer VisionSparsity
WZ

Wanlei Zhou

Professor, City University of Macau, Macao
Parallel and Distributed SystemsIT SecuritySecurity and PrivacyCyber Security
TZ

Tianqing Zhu

City University of Macau
PrivacyCyber SecurityMachine LearningAI Security