psifx - Psychological and Social Interactions Feature Extraction Package

📅 2024-07-14
🏛️ arXiv.org
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In psychology and social sciences, manual annotation of multimodal behavioral data is costly, time-consuming, and suffers from low inter-annotator consistency. To address this, we present an open-source, modular, task-oriented multimodal behavioral feature extraction toolkit. The toolkit integrates speaker diarization, automatic speech recognition (ASR), machine translation, and vision-based pose estimation (e.g., MediaPipe and OpenPose) into an end-to-end configurable pipeline, enabling automated extraction of speaker identity, transcribed and translated speech, body/gesture/face pose, and gaze trajectories from audiovisual recordings. Its novel extensible architecture significantly lowers the barrier to entry for non-AI-expert researchers, promoting standardization and democratization of behavioral analysis tools. Empirical evaluation demonstrates high accuracy and plug-and-play usability, substantially improving annotation efficiency and consistency. The toolkit has already enabled multiple real-time studies on dynamic human behavior analysis.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Web Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web dataSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
psifx is a plug-and-play multi-modal feature extraction toolkit, aiming to facilitate and democratize the use of state-of-the-art machine learning techniques for human sciences research. It is motivated by a need (a) to automate and standardize data annotation processes, otherwise involving expensive, lengthy, and inconsistent human labor, such as the transcription or coding of behavior changes from audio and video sources; (b) to develop and distribute open-source community-driven psychology research software; and (c) to enable large-scale access and ease of use to non-expert users. The framework contains an array of tools for tasks, such as speaker diarization, closed-caption transcription and translation from audio, as well as body, hand, and facial pose estimation and gaze tracking from video. The package has been designed with a modular and task-oriented approach, enabling the community to add or update new tools easily. We strongly hope that this package will provide psychologists a simple and practical solution for efficiently a range of audio, linguistic, and visual features from audio and video, thereby creating new opportunities for in-depth study of real-time behavioral phenomena.
Problem

Research questions and friction points this paper is trying to address.

Automate and standardize human labor-intensive data annotation
Develop open-source community-driven psychology research software
Enable large-scale access for non-expert users
Innovation

Methods, ideas, or system contributions that make the work stand out.

Plug-and-play multi-modal feature extraction toolkit
Automates speaker diarization and pose estimation
Modular design for easy community updates
🔎 Similar Papers
No similar papers found.