🤖 AI Summary
This study addresses the interpretable assessment of suicide risk on social media, which requires the simultaneous execution of risk stratification, evidence extraction, and factor identification. Building upon Qwen2.5-Instruct, this work proposes a multi-task joint training strategy employing QLoRA fine-tuning with a masked causal objective, alongside result aggregation mechanisms such as cross-fold consensus and rate matching. The research demonstrates the effectiveness of multi-task optimization and probability averaging in scenarios involving complementary errors. The proposed system achieves an overall score of 0.7738, reaching 0.8089 on Task 1 and 0.6919 on Task 2, thereby validating the advantages of the customized training and aggregation strategies for comprehensive suicide risk evaluation.
📝 Abstract
Explainable suicide-risk assessment requires models not only to estimate risk severity, but also to identify supporting language and the risk and protective factors expressed in a post. We present our system for the IEEE BigData 2026 Cup on Explainable Suicide Risk Assessment on Social Media, which addresses three tasks: risk-level classification, evidence phrase extraction, and multi-label factor identification. Our approach adapts Qwen2.5-Instruct models using quantized low-rank adaptation (QLoRA) and an answer-masked causal language-model objective. We jointly train across all three tasks for risk classification, jointly train on Tasks~1a and 1b for evidence extraction, and adapt Task~2 separately for factor identification. We also tailor aggregation to each output: we average risk-level probabilities from the 32B and 72B models, combine evidence phrases through cross-fold consensus, and calibrate factor-specific decisions through rate matching based on out-of-fold operating points. On the official leaderboard, the final system achieved a composite score of 0.7738, with 0.8089 on Task~1 and 0.6919 on Task~2. Across the evaluated configurations, three-task training performed best for Task~1a, joint training on Tasks~1a and 1b performed best for Task~1b, and task-specific training performed best for Task~2. Probability averaging further improved Task~1a when component models had complementary errors. These findings highlight the value of tailoring both training objectives and aggregation strategies to the output structure of each task within a unified language-model framework.