π€ AI Summary
This study addresses the challenge of semantic perturbations in safe biomedical natural language inference (NLI) within clinical trial settings, where existing models exhibit insufficient robustness. Building upon the SOLAR Instruct large language model, this work proposes a fine-tuning-free approach integrating input manipulation with customized prompting strategies. Specifically, differentiated prompt templates are designed according to the characteristics of individual sections in clinical trial registrations (CTRs), enabling efficient clinical NLI under both zero-shot and few-shot settings. The proposed method achieves a consistency score of 0.72, ranking 14th on the leaderboard. Furthermore, error analysis reveals the limitations of large language models relying on heuristic shortcuts, providing empirical evidence for enhancing semantic understanding capabilities in medical texts.
π Abstract
This paper describes the approach of the UniBuc team in tackling the SemEval 2024 Task 2: Safe Biomedical Natural Language Inference for Clinical Trials. We used SOLAR Instruct, without any fine-tuning, while focusing on input manipulation and tailored prompting. By customizing prompts for individual CTR sections, in both zero-shot and few-shots settings, we managed to achieve a consistency score of 0.72, ranking 14th in the leaderboard. Our thorough error analysis revealed that our model has a tendency to take shortcuts and rely on simple heuristics, especially when dealing with semantic-preserving changes.