EAAI Journal 2026 Journal Article
Coconut germination precise prediction via multimodal fusion with Co-attention networks: A non-destructive precision agriculture and food engineering solution
- Anum Mehmood
- Zemin Wu
- Yu Zhang
- Xinpeng Bai
- Chengxu Sun
- Uzair Aslam Bhatti
- Mengxing Huang
- Shenghuang Lin
Coconut is the fruit of the coconut palm. Due to its characteristics of a long growth cycle and low germination rate, accurate prediction of its developmental status is particularly important. Traditional research primarily relies on the sectioning method to observe its internal structure. Although this approach can reveal morphological characteristics, its destructive nature prevents continuous monitoring of the internal developmental processes. Recent advancements in Computed Tomography (CT)-based nondestructive imaging and artificial intelligence have enabled novel approaches for investigating internal coconut morphology. However, current methodologies frequently overlook the impact of field environmental factors on coconut germination processes, consequently constraining prediction accuracy. To address this issue, this study proposes a Transformer-based multimodal feature fusion predictive model. Through the integration of CT images and environmental data, the model achieves precise prediction of coconut developmental status. Initially, the enhanced Deeplab V3+ model extracts deep semantic features from coconut CT images, while Fourier positional encoding is applied to amplify periodic features in environmental data (e. g. , temperature, humidity). Subsequently, a cross-modal multi-head attention mechanism is designed to achieve comprehensive fusion between CT-derived semantic features and field data characteristics, thoroughly exploring their correlations. Ultimately, to further enhance model performance, this study incorporates supervised contrastive loss functions and implements intra-class feature aggregation coupled with inter-class feature separation strategies for feature space optimization. The experimental results demonstrate that the proposed model achieves superior performance in coconut developmental stage prediction tasks: compared with conventional unimodal approaches, it improves prediction accuracy and F1-score by 9 % and 8 %, respectively, thereby validating the effectiveness of multimodal data fusion and the rationality of the model design.