THPL: A Vision-to-Language Decision Support Framework for Rainbow Trout Feeding Management in RAS

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the misalignment between fish behavior and management knowledge in recirculating aquaculture systems, which renders feeding decisions difficult to execute and interpret. To this end, we propose a vision-to-language generative decision-making framework. Specifically, an activity coefficient is constructed via trajectory extraction, while a hierarchical encoder fuses kinematic and dual-evidence representations. A large language model is then fine-tuned using LoRA with environmental parameters to perform causal reasoning. Furthermore, counterfactual multimodal Direct Preference Optimization (mDPO) is introduced to mitigate template bias and enhance the causal consistency of decisions. Experimental results demonstrate that the activity coefficient correlates significantly with expert annotations. Decision accuracy improves substantially from 33.33% to 96.67%, achieving a METEOR score of 85.30%, alongside markedly enhanced textual diversity.
📝 Abstract
In Recirculating Aquaculture Systems (RAS), precision feeding is critical for minimizing costs and improving fish welfare. However, existing methods lack cognitive alignment between fish behaviors and management knowledge, impeding translation into executable, interpretable feeding decisions. To address this, we propose THPL, a generative feeding decision framework tailored for rainbow trout (Oncorhynchus mykiss) in RAS. First, Fishsort extracts trajectories to establish an Activity Coefficient (AC) quantifying feeding intensity. Second, a Hierarchical Behavior Encoder (HBE) models individual temporal progression and collective dynamics using Temporal and Set Transformers, transforming trajectory tensors into dual-evidence representations of explicit physical and implicit soft tokens. Finally, these tokens are integrated with environmental parameters, metadata, and expert rules to fine-tune an LLM via LoRA, followed by counterfactual multimodal Direct Preference Optimization (mDPO) to reinforce causal reasoning. Results show that AC exhibits a statistically significant monotonic positive correlation with expert-annotated feeding intensity (Spearman $ρ= 0.925$, $p < 0.001$). Ablations indicate that decision accuracy improves from 33.33% (text-only baseline) to 93.33% with dual-evidence tokens, confirming that continuous spatiotemporal tokens provide necessary physical grounding for LLMs. Compared with standard LoRA, counterfactual mDPO elevates decision accuracy from 93.33% to 96.67%, advances METEOR from 58.10% to 85.30%, reduces Self-BLEU-2 from 58.79% to 52.88%, and increases Distinct-3 from 6.68% to 7.81%, suppressing templating and actuation biases while reinforcing causal consistency and operational safety. Overall, by integrating continuous kinematics with LLM reasoning, this study provides a novel decision support paradigm for precision aquaculture.
Problem

Research questions and friction points this paper is trying to address.

precision feeding
recirculating aquaculture systems
cognitive alignment
feeding decision
fish behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-to-Language Framework
Hierarchical Behavior Encoder
Dual-evidence Tokens
Counterfactual Multimodal DPO
Precision Aquaculture
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Meng Liang
College of Biosystems Engineering and Food Science, Zhejiang University, Hangzhou 310058, Zhejiang, PR China
G
Guanbo Feng
College of Biosystems Engineering and Food Science, Zhejiang University, Hangzhou 310058, Zhejiang, PR China
H
Haozhuang Chi
School of Mechanical and Aerospace Engineering, Nanyang Technological University, Singapore 639798, Singapore
Shilong Zhao
Shilong Zhao
University of Chinese Academy of Sciences
Z
Zhixin Xiong
College of Engineering, Nanjing Agricultural University, Nanjing 210031, Jiangsu, PR China
Yuhang He
Yuhang He
Microsoft Research
Multimodal LearningMachine LearningWorld ModelComputer VisionSpatial Audio
W
Wenfeng Han
College of Biosystems Engineering and Food Science, Zhejiang University, Hangzhou 310058, Zhejiang, PR China
T
Tianhao Zhao
College of Biosystems Engineering and Food Science, Zhejiang University, Hangzhou 310058, Zhejiang, PR China
Z
Zhihong Ma
College of Biosystems Engineering and Food Science, Zhejiang University, Hangzhou 310058, Zhejiang, PR China
Y
Ying Liu
College of Biosystems Engineering and Food Science, Zhejiang University, Hangzhou 310058, Zhejiang, PR China