Arrow Research search

Author name cluster

Zhihao Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

15 papers
2 author rows

Possible papers

15

EAAI Journal 2026 Journal Article

Biaxial tensile test meets unsupervised learning: a novel strategy for accurate constitutive modeling

  • Zhihao Wang
  • Dominique Guines
  • Xingrong Chu
  • Lionel Leotoing
  • Wenke Wang

With the rapid advancements in forming processes and materials engineering, the demand for precise constitutive modeling has grown significantly, and machine learning techniques have shown great potential in this field, offering more flexible and efficient solutions. This study proposes a new unsupervised learning approach combining an artificial neural network (ANN) with the finite element (FE) method to learn thermo-viscoplastic constitutive relations from indirect data. The ANN maps the strain, temperature and strain rate to flow stress of AA6061 sheets and is implemented into finite element analysis via a user material subroutine that adopts the Yld2000-3d yield surface and associated flow rule assumptions. The unsupervised learning approach eliminates the need for flow stress computation, relying instead on strain and global force measurements provided by fifteen biaxial tensile tests that use a dedicated cruciform specimen at different temperatures and strain rates. These indirect data contain rich information about material constitutive relations and enable ANN constitutive modeling in complex scenarios. A novel symbolic differentiation approximation method is developed to address the challenge of the gradient computation of a loss function that contains strains from FE solutions. The method is applicable to general FE framework, offers robust and efficient convergence, and is user-friendly to implement. Additionally, a continual training strategy with weighted gradients by gradually introducing new dataset is adopted to balance multiple datasets and mitigate catastrophic forgetting. Finally, validations through thermal equi-biaxial stretching and shear tests demonstrate that the trained ANN model can accurately predict constitutive relations and reasonably extrapolate to large strain levels.

AAAI Conference 2026 Conference Paper

CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution

  • Baoliang Tian
  • Yuxuan Si
  • Jilong Wang
  • LingYao Li
  • Zhongyuan Bao
  • Zineng Zhou
  • Tao Wang
  • Sixu Li

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual cues often conflict, requiring models to perform structured reasoning beyond surface-level alignment. We introduce CrossCheck-Bench, a diagnostic benchmark for evaluating contradiction detection in multimodal inputs. The benchmark adopts a hierarchical task framework covering three levels of reasoning complexity and defines seven atomic capabilities essential for resolving cross-modal inconsistencies. CrossCheck-Bench includes 15k question-answer pairs sourced from real-world artifacts with synthetically injected contradictions. The dataset is constructed through a multi-stage annotation pipeline involving more than 450 expert hours to ensure semantic validity and calibrated difficulty across perception, integration, and reasoning. We evaluate 13 state-of-the-art vision-language models and observe a consistent performance drop as tasks shift from perceptual matching to logical contradiction detection. Most models perform well on isolated entity recognition but fail when multiple clues must be synthesized for conflict reasoning. Capability-level analysis further reveals uneven skill acquisition, especially in tasks requiring multi-step inference or rule-based validation. Additional probing shows that conventional prompting strategies such as Chain-of-Thought and Set-of-Mark yield only marginal gains. By contrast, methods that interleave symbolic reasoning with grounded visual processing achieve more stable improvements. These results highlight a persistent bottleneck in multimodal reasoning and suggest new directions for building models capable of robust cross-modal verification.

AAAI Conference 2026 Conference Paper

EcoDiffusion: Uncertainty-Aware Emulation of Ecosystem Processes with Conditional Diffusion for Long Sequences with Single-Step Initialization

  • Ruohan Li
  • Zhihao Wang
  • Xiaowei Jia
  • Gengchen Mai
  • Lei Ma
  • George C. Hurtt
  • Quan Shen
  • Zhili Li

Terrestrial ecosystems constitute a major component of the global carbon sink and play a critical role in regulating the global carbon cycle. Although process-based models such as the Ecosystem Demography (ED) model are widely used to simulate these dynamics and widely adopted in research and applications, they remain computationally intensive and are not well suited for large-scale (e.g., global) projections at high spatial and temporal resolution, or under wide-range of future scenarios. AI-based emulators of process-based physical models have emerged as promising ways to accelerate the computation. However, there are several challenges in developing emulators for ecosystem processes, including error accumulation over long sequences, single-step initial conditions, and high-dimensional environmental conditions. Existing works often rely on time-series patterns in look-back windows, which are not well-suited for the problem with single-step initial conditions. Moreover, they often do not consider uncertainty, making it hard to know when the approximations are highly confident and when the results may need to be updated, e.g., by the process-based models. To address these limitations, we introduce EcoDiffusion, a conditional diffusion framework tailored for ecosystem dynamics emulation. We evaluated EcoDiffusion at locations distributed worldwide under different scenarios and showed that it demonstrated significant improvements over existing models.

AAAI Conference 2026 Conference Paper

Federated Context-Aware Personalized Recommendation

  • Zhihao Wang
  • Xiaoying Liao
  • Wenke Huang
  • Bingqian Liu
  • Tian Chen
  • Jian Wang
  • Bing Li

Federated recommender system is emerging as a new paradigm for providing personalized services while preserving user data privacy. Most existing personalized federated recommender systems predict the user's next item by discretely training user and item embeddings. However, this training approach overlooks the user's behavioral patterns, suffers from low interpretability, and requires a substantial amount of data and meticulous fine-tuning to achieve stable and accurate embeddings. To address these limitations, we propose Federated Context-Aware Personalized Recommendation (FedCAR), a novel framework that leverages users’ recent interactions as behavioral context to guide prediction. Instead of static user embeddings, FedCAR dynamically constructs context representations by aggregating and weighting recently interacted item embeddings. Additionally, we incorporate a contrastive learning strategy that enables the model to capture shared behavioral structures across clients while maintaining personalized preferences, enhancing both generalization and robustness in heterogeneous settings. Experiments on 5 benchmark datasets show that FedCAR consistently outperforms state-of-the-art methods and provides interpretable recommendations by explicitly modeling context dependencies.

IJCAI Conference 2025 Conference Paper

An Empirical Study of Federated Prompt Learning for Vision Language Model

  • Zhihao Wang
  • Wenke Huang
  • Tian Chen
  • Zekun Shi
  • Guancheng Wan
  • Yu Qiao
  • Bin Yang
  • Jian Wang

The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt learning with VLM in federated learning (FL) scenarios remains underexplored. This paper systematically investigates the behavioral differences between language prompt learning (LPT) and vision prompt learning (VPT) under data heterogeneity challenges, including label skew and domain shift. We conduct extensive experiments to evaluate the impact of various FL and prompt configurations, such as client scale, aggregation strategies, and prompt length, to assess the robustness of Federated Prompt Learning (FPL). Furthermore, we explore strategies for enhancing prompt learning in complex scenarios where label skew and domain shift coexist, including leveraging both prompt types when computational resources allow. Our findings offer practical insights into optimizing prompt learning in federated settings, contributing to the broader deployment of VLMs in privacy-preserving environments.

NeurIPS Conference 2025 Conference Paper

CarbonGlobe: A Global-Scale, Multi-Decade Dataset and Benchmark for Carbon Forecasting in Forest Ecosystems

  • Zhihao Wang
  • Lei Ma
  • George Hurtt
  • Xiaowei Jia
  • Yanhua Li
  • Ruohan Li
  • Zhili Li
  • Shuo Xu

Forest ecosystems play a critical role in the Earth system as major carbon sinks that are essential for carbon neutralization and climate change mitigation. However, the Earth has undergone significant deforestation and forest degradation, and the remaining forested areas are also facing increasing pressures from socioeconomic factors and climate change, potentially pushing them towards tipping points. Responding to the grand challenge, a theory-based Ecosystem Demography (ED) model has been continuously developed over the past two decades and serves as a key component in major initiatives, including the Global Carbon Budget, NASA Carbon Monitoring System, and US Greenhouse Gas Center. Despite its growing importance in combating climate change and shaping carbon policies, ED's expensive computation significantly limits its ability to estimate carbon dynamics at the global scale with high spatial resolution. Recently, machine learning (ML) models have shown promising potential in approximating theory-based models with interesting success in various domains including weather forecasting, thanks to the open-source benchmark datasets made available. However, there are currently no publicly available ML-ready datasets for global carbon dynamics forecasting in forest ecosystems. The limited data availability hinders the development of corresponding ML emulators. Furthermore, the inputs needed for running ED are highly complex with over a hundred variables from various remote sensing products. To bridge the gap, we develop a new ML-ready benchmark dataset, \textit{CarbonGlobe}, for carbon dynamics forecasting, featuring that: (1) the data has a global-scale coverage at 0. 5$^\circ$ resolution; (2) the temporal range spans 40 years; (3) the inputs integrate extensive multi-source data from different sensing products, with calibrated outputs from ED; (4) the data is formatted in ML-ready forms and split into different evaluation scenarios based on climate conditions, etc. ; (5) a set of problem-driven metrics is designed to develop benchmarks using various ML models to best align with the needs of downstream applications. Our dataset and code are publicly available on Kaggle and GitHub: https: //www. kaggle. com/datasets/zhihaow/carbonglobe and https: //github. com/zhwang0/carbon-globe.

ICML Conference 2025 Conference Paper

Efficient Robotic Policy Learning via Latent Space Backward Planning

  • Dongxiu Liu
  • Haoyi Niu
  • Zhihao Wang
  • Jinliang Zheng
  • Yinan Zheng
  • Zhonghong Ou
  • Jianming Hu
  • Jianxiong Li

Current robotic planning methods often rely on predicting multi-frame images with full pixel details. While this fine-grained approach can serve as a generic world model, it introduces two significant challenges for downstream policy learning: substantial computational costs that hinder real-time deployment, and accumulated inaccuracies that can mislead action extraction. Planning with coarse-grained subgoals partially alleviates efficiency issues. However, their forward planning schemes can still result in off-task predictions due to accumulation errors, leading to misalignment with long-term goals. This raises a critical question: Can robotic planning be both efficient and accurate enough for real-time control in long-horizon, multi-stage tasks? To address this, we propose a B ackward P lanning scheme in L atent space ( LBP ), which begins by grounding the task into final latent goals, followed by recursively predicting intermediate subgoals closer to the current state. The grounded final goal enables backward subgoal planning to always remain aware of task completion, facilitating on-task prediction along the entire planning horizon. The subgoal-conditioned policy incorporates a learnable token to summarize the subgoal sequences and determines how each subgoal guides action extraction. Through extensive simulation and real-robot long-horizon experiments, we show that LBP outperforms existing fine-grained and forward planning methods, achieving SOTA performance. Project Page: https: //lbp-authors. github. io.

AAAI Conference 2025 Conference Paper

Federated Recommendation with Explicitly Encoding Item Bias

  • Zhihao Wang
  • He Bai
  • Wenke Huang
  • Duantengchuan Li
  • Jian Wang
  • Bing Li

With the development of federated learning techniques and the increased need for user privacy protection, the federated recommendation has become a new recommendation paradigm. However, most existing works focus on user-level federated recommendation, leaving platform-level federated recommendation largely unexplored. A significant challenge in platform-level federated recommendation scenarios is severe label skew. Users behave in various ways on different platforms, bringing up the rating and item bias problem. In this work, we propose FREIB (Federated Recommendation with Explicitly Encoding Item Bias). The core idea is explicitly encoding item bias during federated learning, addressing the problem of fuzzy item bias, and achieving consistent representation in label skew scenarios. We achieve this by utilizing global knowledge guidance to model common rating patterns and by aligning feature prototypes to enhance item encoding at the same rating level. Extensive experiments conducted on three public datasets demonstrate the superiority of our method over several state-of-the-art approaches.

IJCAI Conference 2025 Conference Paper

Pixel-wise Divide and Conquer for Federated Vessel Segmentation

  • Tian Chen
  • Wenke Huang
  • Zhihao Wang
  • Zekun Shi
  • He Li
  • Wenhui Dong
  • Mang Ye
  • Bo Du

Accurate vessel segmentation is essential for diagnosing and managing vascular and ophthalmic diseases. Traditional learning-based vessel segmentation methods heavily rely on high-quality, pixel-level annotated datasets. However, segmentation performance suffers significantly when applied in federated learning settings due to vessel morphology inconsistency and vessel-background imbalance. The former limits the ability of models to capture fine-grained vessels, while the latter overemphasizes background pixels and biases the model towards them. To address these challenges, we propose a novel method named Federated Vessel-Aware Calibration (FVAC), which leverages global uncertainty to provide differentiated guidance for clients, focusing on pixels of various morphologies that are difficult to distinguish. Furthermore, we introduce a foreground-background decoupling alignment strategy that utilizes more stable and balanced global features to mitigate semantic drift caused by vessel-background imbalance in local clients. Comprehensive experiments confirm the effectiveness of our method

ICRA Conference 2025 Conference Paper

Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning

  • Jianxiong Li
  • Zhihao Wang
  • Jinliang Zheng
  • Xiaoai Zhou
  • Guanming Wang
  • Guanglu Song
  • Yu Liu 0015
  • Jingjing Liu

Multimodal task specification is essential for enhanced robotic performance, where Cross-modality Alignment enables the robot to holistically understand complex task instructions. Directly annotating multimodal instructions for model training proves impractical, due to the sparsity of paired multimodal data. In this study, we demonstrate that by leveraging unimodal instructions abundant in real data, we can effectively teach robots to learn multimodal task specifications. First, we endow the robot with strong Crossmodality Alignment capabilities, by pretraining a robotic multimodal encoder using extensive out-of-domain data. Then, we employ two Collapse and Corrupt operations to further bridge the remaining modality gap in the learned multimodal representation. This approach projects different modalities of identical task goal as interchangeable representations, thus enabling accurate robotic operations within a well-aligned multimodal latent space. Evaluation across more than 130 tasks and 4000 evaluations on both simulated LIBERO benchmark and real robot platforms showcases the superior capabilities of our proposed framework, demonstrating significant potential in overcoming data constraints in robotic learning. Website: zh1hao. wang/Robo_MUTUAL

NeurIPS Conference 2025 Conference Paper

TreeFinder: A US-Scale Benchmark Dataset for Individual Tree Mortality Monitoring Using High-Resolution Aerial Imagery

  • Zhihao Wang
  • Cooper Li
  • Ruichen Wang
  • Lei Ma
  • George Hurtt
  • Xiaowei Jia
  • Gengchen Mai
  • Zhili Li

Monitoring individual tree mortality at scale has been found to be crucial for understanding forest loss, ecosystem resilience, carbon fluxes, and climate-induced impacts. However, the fine-granularity monitoring faces major challenges on both the data and methodology sides because: (1) finding isolated individual-level tree deaths requires high-resolution remote sensing images with broad coverage, and (2) compared to regular geo-objects (e. g. , buildings), dead trees often exhibit weaker contrast and high variability across tree types, landscapes and ecosystems. Existing datasets on tree mortality primarily rely on moderate-resolution satellite imagery (e. g. , 30m resolution), which aims to detect large-patch wipe-outs but is unable to recognize individual-level tree mortality events. Several efforts have explored alternatives via very-high-resolution drone imagery. However, drone images are highly expensive and can only be collected at local scales, which are therefore not suitable for national-scale applications and beyond. To bridge the gaps, we introduce TreeFinder, the first high-resolution remote sensing benchmark dataset designed for individual-level tree mortality mapping across the Contiguous United States (CONUS). Specifically, the dataset uses NAIP imagery at 0. 6m resolution that provides wall-to-wall coverage of the entire CONUS. TreeFinder contains images with pixel-level labels generated via extensive manual annotation that covers forested areas in 48 states with over 23, 000 hectares. All annotations are rigorously validated using multi-temporal NAIP images and auxiliary vegetation indices from remote sensing imagery. Moreover, TreeFinder includes multiple evaluation scenarios to test the models' ability in generalizing across different geographic regions, climate zones, and forests with different plant function types. Finally, we develop benchmarks using a suite of semantic segmentation models, including both convolutional architectures and more recent foundation models based on vision transformers for general and remote sensing images. Our dataset and code are publicly available on Kaggle and GitHub: https: //www. kaggle. com/datasets/zhihaow/tree-finder and https: //github. com/zhwang0/treefinder.

AAAI Conference 2024 Conference Paper

Existence Is Chaos: Enhancing 3D Human Motion Prediction with Uncertainty Consideration

  • Zhihao Wang
  • Yulin Zhou
  • Ningyu Zhang
  • Xiaosong Yang
  • Jun Xiao
  • Zhao Wang

Human motion prediction is consisting in forecasting future body poses from historically observed sequences. It is a longstanding challenge due to motion's complex dynamics and uncertainty. Existing methods focus on building up complicated neural networks to model the motion dynamics. The predicted results are required to be strictly similar to the training samples with L2 loss in current training pipeline. However, little attention has been paid to the uncertainty property which is crucial to the prediction task. We argue that the recorded motion in training data could be an observation of possible future, rather than a predetermined result. In addition, existing works calculate the predicted error on each future frame equally during training, while recent work indicated that different frames could play different roles. In this work, a novel computationally efficient encoder-decoder model with uncertainty consideration is proposed, which could learn proper characteristics for future frames by a dynamic function. Experimental results on benchmark datasets demonstrate that our uncertainty consideration approach has obvious advantages both in quantity and quality. Moreover, the proposed method could produce motion sequences with much better quality that avoids the intractable shaking artefacts. We believe our work could provide a novel perspective to consider the uncertainty quality for the general motion prediction task and encourage the studies in this field. The code will be available in https://github.com/Motionpre/Adaptive-Salient-Loss-SAGGB.

AAAI Conference 2024 Conference Paper

SimFair: Physics-Guided Fairness-Aware Learning with Simulation Models

  • Zhihao Wang
  • Yiqun Xie
  • Zhili Li
  • Xiaowei Jia
  • Zhe Jiang
  • Aolin Jia
  • Shuo Xu

Fairness-awareness has emerged as an essential building block for the responsible use of artificial intelligence in real applications. In many cases, inequity in performance is due to the change in distribution over different regions. While techniques have been developed to improve the transferability of fairness, a solution to the problem is not always feasible with no samples from the new regions, which is a bottleneck for pure data-driven attempts. Fortunately, physics-based mechanistic models have been studied for many problems with major social impacts. We propose SimFair, a physics-guided fairness-aware learning framework, which bridges the data limitation by integrating physical-rule-based simulation and inverse modeling into the training design. Using temperature prediction as an example, we demonstrate the effectiveness of the proposed SimFair in fairness preservation.

NeurIPS Conference 2024 Conference Paper

SolarCube: An Integrative Benchmark Dataset Harnessing Satellite and In-situ Observations for Large-scale Solar Energy Forecasting

  • Ruohan Li
  • Yiqun Xie
  • Xiaowei Jia
  • Dongdong Wang
  • Yanhua Li
  • Yingxue Zhang
  • Zhihao Wang
  • Zhili Li

Solar power is a critical source of renewable energy, offering significant potential to lower greenhouse gas emissions and mitigate climate change. However, the cloud induced-variability of solar radiation reaching the earth’s surface presents a challenge for integrating solar power into the grid (e. g. , storage and backup management). The new generation of geostationary satellites such as GOES-16 has become an important data source for large-scale and high temporal frequency solar radiation forecasting. However, no machine-learning-ready dataset has integrated geostationary satellite data with fine-grained solar radiation information to support forecasting model development and benchmarking with consistent metrics. We present SolarCube, a new ML-ready benchmark dataset for solar radiation forecasting. SolarCube covers 19 study areas distributed over multiple continents: North America, South America, Asia, and Oceania. The dataset supports short (i. e. , 30 minutes to 6 hours) and long-term (i. e. , day-ahead or longer) solar radiation forecasting at both point-level (i. e. , specific locations of monitoring stations) and area-level, by processing and integrating data from multiple sources, including geostationary satellite images, physics-derived solar radiation, and ground station observations from different monitoring networks over the globe. We also evaluated a set of forecasting models for point- and image-based time-series data to develop performance benchmarks under different testing scenarios. The dataset is available at https: //doi. org/10. 5281/zenodo. 11498739. A Python library is available to conveniently generate different variations of the dataset based on user needs, along with baseline models at https: //github. com/Ruohan-Li/SolarCube.

YNIMG Journal 2023 Journal Article

The domain-separation language network dynamics in resting state support its flexible functional segregation and integration during language and speech processing

  • Binke Yuan
  • Hui Xie
  • Zhihao Wang
  • Yangwen Xu
  • Hanqing Zhang
  • Jiaxuan Liu
  • Lifeng Chen
  • Chaoqun Li

Modern linguistic theories and network science propose that language and speech processing are organized into hierarchical, segregated large-scale subnetworks, with a core of dorsal (phonological) stream and ventral (semantic) stream. The two streams are asymmetrically recruited in receptive and expressive language or speech tasks, which showed flexible functional segregation and integration. We hypothesized that the functional segregation of the two streams was supported by the underlying network segregation. A dynamic conditional correlation approach was employed to construct framewise time-varying language networks and k-means clustering was employed to investigate the temporal-reoccurring patterns. We found that the framewise language network dynamics in resting state were robustly clustered into four states, which dynamically reconfigured following a domain-separation manner. Spatially, the hub distributions of the first three states highly resembled the neurobiology of speech perception and lexical-phonological processing, speech production, and semantic processing, respectively. The fourth state was characterized by the weakest functional connectivity and was regarded as a baseline state. Temporally, the first three states appeared exclusively in limited time bins (∼15%), and most of the time (> 55%), state 4 was dominant. Machine learning-based dFC-linguistics prediction analyses showed that dFCs of the four states significantly predicted individual linguistic performance. These findings suggest a domain-separation manner of language network dynamics in resting state, which forms a dynamic "meta-network" framework to support flexible functional segregation and integration during language and speech processing.

v2026.09.13