Arrow Research search

Author name cluster

Jiawei Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
2 author rows

Possible papers

10

EAAI Journal 2026 Journal Article

A surrogate model for single-stage synchronous induction coilgun based on scientific machine learning

  • Jialu Gao
  • Gangquan Si
  • Xikui Ma
  • Jiawei Wang
  • Yating Wang

The synchronous induction coilgun is a representative electromechanical coupled system whose numerical modeling usually relies on the finite element method (FEM) and the current filament method (CFM). However, these physics-based approaches are often computationally intensive and time-consuming in design and optimization processes. This study aims to develop a data-efficient operator-learning framework, termed Gated Recurrent Unit and Fully Connected Operator Network (G&F-ONet), for predicting the velocity trajectory of a single-stage induction coilgun. Built upon the Deep Operator Network (DeepONet) architecture, the proposed model embeds physical prior knowledge by introducing a recurrent branch whose structure reflects the evolution pattern of geometric parameters during launch. The dual-branch vanilla Multiple-Input Operator Network (MIONet) is used as the baseline, and the relative L 2 error serves as the primary evaluation metric. Experimental results show that, using a dataset generated through Latin Hypercube Sampling (LHS) from a CFM-based model, G&F-ONet reduces the average relative L 2 error by 43. 1% compared with the baseline model and exhibits superior data efficiency, robustness, and generalization capability. All experiments are performed on an NVIDIA RTX 2060 Graphics Processing Unit (GPU) with an Intel i5 Central Processing Unit (CPU). The results demonstrate that the proposed approach, which introduces physical priors through inductive bias, can effectively learn the complex mapping between coilgun parameters and launch velocity, achieving fast, reliable, and physically interpretable predictions. Overall, the proposed framework provides an efficient surrogate model for the design and optimization of synchronous induction coilguns.

EAAI Journal 2025 Journal Article

A unified solution for replacing position embedding in Vision Transformer for object detection

  • Xuezhuan Zhao
  • Jiawei Wang
  • Lingling Li
  • XiaoYan Shao
  • Kexin Zhang

The traditional Vision Transformer (ViT) demonstrating outstanding performance in various computer vision tasks. However, this high performance relies on substantial training on large datasets, for position embedding which is the key part of ViT, requires extensive data to maximize their effectiveness. What is more, in diverse vision task scenarios, position embedding may not perform well and even yield redundant information, which limits the overall model’s flexibility and robustness. To tackle this problem, we propose a novel position-free embedding model HV-SwinViT, which embed both Horizontal and Vertical features by using the self-attention mechanism with the position-free functions, replacing the positional embedding. To achieve this, firstly, we propose a module Horizontal and Vertical orientation Swin Transformer block ( HVblock ), which generate feature maps containing both horizontal and vertical information by adopting fully connected layers, then we design two hybrid sub-network: HVblock in Backbone ( HV-Swin-B ) and Channel Fusion 2 with HVblock ( Cf2-HV ) which leverage convolution layer and HVblock to tackle positional relationship information. Also we defined a learnable nonlinear activation function to increase the sensitivity to nonlinear position features. Experimental results demonstrate that the proposed HV-SwinViT model achieves an improvement of over 0. 1% in Average Precision (AP) compared to state-of-the-art methods on the MS-COCO2017 dataset. Additionally, our model outperforms several network architectures designed specifically for aerial photography of targets, attaining an A P score of 35. 6% and 32. 1% on the VisDrone2019 and AI-TOD, respectively. These results confirm that the HV-SwinViT model can be a unified solution and highlight the stability and robustness of our approach across diverse scenarios.

IJCAI Conference 2025 Conference Paper

DGCPL: Dual Graph Distillation for Concept Prerequisite Relation Learning

  • Miao Zhang
  • Jiawei Wang
  • Jinying Han
  • Kui Xiao
  • Zhifei Li
  • Yan Zhang
  • Hao Chen
  • Shihui Wang

Concept prerequisite relations determine the learning order of knowledge concepts in one domain, which has an important impact on teachers' course design and students' personalized learning. Current research usually predicts concept prerequisite relations from the perspective of knowledge, and rarely pays attention to the role of learners' learning behavior. We propose a Dual Graph Distillation Method for Concept Prerequisite Relation Learning (DGCPL). Specifically, DGCPL constructs a dual graph structure from both the knowledge and learning behavior perspectives, and captures the high-order knowledge features and learning behavior features through the concept-resource hypergraph and the learning behavior graph respectively. In addition, we introduce a gated knowledge distillation to fuse the structural information of concept nodes in the two graphs, so as to obtain a more comprehensive concept embedding representation and achieve accurate prediction of prerequisite relations. On three public benchmark datasets, we compare DGCPL with eight graph-based baseline methods and five traditional classification baseline methods. The experimental results show that DGCPL achieves state-of-the-art performance in learning concept prerequisite relations. Our code is available at https: //github. com/wisejw/DGCPL.

AAAI Conference 2025 Conference Paper

Learning Concept Prerequisite Relation via Global Knowledge Relation Optimization

  • Miao Zhang
  • Jiawei Wang
  • Kui Xiao
  • Shihui Wang
  • Yan Zhang
  • Hao Chen
  • Zhifei Li

Learning concept prerequisite relations helps better master and build a logically coherent knowledge structure. Many studies use graph neural networks to create heterogeneous knowledge networks that enhance concept representations. However, different types of relations in these networks can influence each other. Existing research often focuses solely on concept relations, neglecting other types of knowledge connections. To address this issue, this paper proposes a novel concept prerequisite relation learning model, named the Global Knowledge Relation Optimization Model(GKROM). Specifically, we capture the impact of different knowledge relation types on document and concept semantic representations separately, integrating the document and concept semantic representations. Then, we introduce multi-objective learning to optimize the knowledge relation network from a global perspective. Through the above optimization, GKROM learns richer semantic representations for concepts and documents, improving the accuracy of concept prerequisite relation learning. Extensive experiments on public datasets demonstrate the effectiveness of our GKROM, achieving state-of-the-art performance in concept prerequisite relation learning.

ICML Conference 2024 Conference Paper

Boximator: Generating Rich and Controllable Motions for Video Synthesis

  • Jiawei Wang
  • Yuchen Zhang
  • Jiaxin Zou
  • Yan Zeng
  • Guoqiang Wei
  • Liping Yuan
  • Hang Li

Generating rich and controllable motion is a pivotal challenge in video synthesis. We propose Boximator, a new approach for fine-grained motion control. Boximator introduces two constraint types: hard box and soft box. Users select objects in the conditional frame using hard boxes and then use either type of boxes to roughly or rigorously define the object’s position, shape, or motion path in future frames. Boximator functions as a plug-in for existing video diffusion models. Its training process preserves the base model’s knowledge by freezing the original weights and training only the control module. To address training challenges, we introduce a novel self-tracking technique that greatly simplifies the learning of box-object correlations. Empirically, Boximator achieves state-of-the-art video quality (FVD) scores, improving on two base models, and further enhanced after incorporating box constraints. Its robust motion controllability is validated by drastic increases in the bounding box alignment metric. Human evaluation also shows that users favor Boximator generation results over the base model.

NeurIPS Conference 2024 Conference Paper

Large Language Models as Urban Residents: An LLM Agent Framework for Personal Mobility Generation

  • Jiawei Wang
  • Renhe Jiang
  • Chuang Yang
  • Zengqing Wu
  • Makoto Onizuka
  • Ryosuke Shibasaki
  • Noboru Koshizuka
  • Chuan Xiao

This paper introduces a novel approach using Large Language Models (LLMs) integrated into an agent framework for flexible and effective personal mobility generation. LLMs overcome the limitations of previous models by effectively processing semantic data and offering versatility in modeling various tasks. Our approach addresses three research questions: aligning LLMs with real-world urban mobility data, developing reliable activity generation strategies, and exploring LLM applications in urban mobility. The key technical contribution is a novel LLM agent framework that accounts for individual activity patterns and motivations, including a self-consistency approach to align LLMs with real-world activity data and a retrieval-augmented strategy for interpretable activity generation. We evaluate our LLM agent framework and compare it with state-of-the-art personal mobility generation approaches, demonstrating the effectiveness of our approach and its potential applications in urban mobility. Overall, this study marks the pioneering work of designing an LLM agent framework for activity generation based on real-world human activity data, offering a promising tool for urban mobility analysis.

ICLR Conference 2023 Conference Paper

Towards Effective and Interpretable Human-Agent Collaboration in MOBA Games: A Communication Perspective

  • Yiming Gao 0007
  • Feiyu Liu
  • Liang Wang 0015
  • Zhenjie Lian
  • Weixuan Wang
  • Siqin Li
  • Xianliang Wang
  • Xianhan Zeng

MOBA games, e.g., Dota2 and Honor of Kings, have been actively used as the testbed for the recent AI research on games, and various AI systems have been developed at the human level so far. However, these AI systems mainly focus on how to compete with humans, less on exploring how to collaborate with humans. To this end, this paper makes the first attempt to investigate human-agent collaboration in MOBA games. In this paper, we propose to enable humans and agents to collaborate through explicit communication by designing an efficient and interpretable Meta-Command Communication-based framework, dubbed MCC, for accomplishing effective human-agent collaboration in MOBA games. The MCC framework consists of two pivotal modules: 1) an interpretable communication protocol, i.e., the Meta-Command, to bridge the communication gap between humans and agents; 2) a meta-command value estimator, i.e., the Meta-Command Selector, to select a valuable meta-command for each agent to achieve effective human-agent collaboration. Experimental results in Honor of Kings demonstrate that MCC agents can collaborate reasonably well with human teammates and even generalize to collaborate with different levels and numbers of human teammates. Videos are available at https://sites.google.com/view/mcc-demo.

ICLR Conference 2023 Conference Paper

Write and Paint: Generative Vision-Language Models are Unified Modal Learners

  • Shizhe Diao
  • Wangchunshu Zhou
  • Xinsong Zhang
  • Jiawei Wang

Recent advances in vision-language pre-training have pushed the state-of-the-art on various vision-language tasks, making machines more capable of multi-modal writing (image-to-text generation) and painting (text-to-image generation). However, few studies investigate if these two essential capabilities can be learned together and boost each other, making a versatile and powerful multi-modal foundation model. In this work, we disclose the potential of symmetric generative vision-language pre-training in learning to write and paint concurrently, and propose a new unified modal model, named DaVinci, trained with prefix language modeling and prefix image modeling, a simple generative self-supervised objective on image-text pairs. Thanks to the proposed prefix multi-modal modeling framework, DaVinci is simple to train, scalable to huge data, adaptable to both writing and painting tasks, and also strong on other vision, text, and multi-modal understanding tasks. DaVinci achieves competitive performance on a wide range of 27 generation/understanding tasks and demonstrates the superiority of combining vision/language generative pre-training. Furthermore, we carefully benchmark the performance of different vision-language pre-training objectives on different scales of pre-training datasets on a heterogeneous and broad distribution coverage. Our results demonstrate the potential of exploiting self-supervision in both language and vision inputs, and establish new, stronger baselines for future comparisons at different data scales. The code and pre-trained models are available at https://github.com/shizhediao/DaVinci.

IJCAI Conference 2021 Conference Paper

Reducing Bus Bunching with Asynchronous Multi-Agent Reinforcement Learning

  • Jiawei Wang
  • Lijun Sun

The bus system is a critical component of sustainable urban transportation. However, due to the significant uncertainties in passenger demand and traffic conditions, bus operation is unstable in nature and bus bunching has become a common phenomenon that undermines the reliability and efficiency of bus services. Despite recent advances in multi-agent reinforcement learning (MARL) on traffic control, little research has focused on bus fleet control due to the tricky asynchronous characteristic---control actions only happen when a bus arrives at a bus stop and thus agents do not act simultaneously. In this study, we formulate route-level bus fleet control as an asynchronous multi-agent reinforcement learning (ASMR) problem and extend the classical actor-critic architecture to handle the asynchronous issue. Specifically, we design a novel critic network to effectively approximate the marginal contribution for other agents, in which graph attention neural network is used to conduct inductive learning for policy evaluation. The critic structure also helps the ego agent optimize its policy more efficiently. We evaluate the proposed framework on real-world bus services and actual passenger demand derived from smart card data. Our results show that the proposed model outperforms both traditional headway-based control methods and existing MARL methods.

AAAI Conference 2013 Conference Paper

Exploring the Contribution of Unlabeled Data in Financial Sentiment Analysis

  • Jimmy Ren
  • Wei Wang
  • Jiawei Wang
  • Stephen Liao

With the proliferation of its applications in various industries, sentiment analysis by using publicly available web data has become an active research area in text classification during these years. It is argued by researchers that semi-supervised learning is an effective approach to this problem since it is capable to mitigate the manual labeling effort which is usually expensive and timeconsuming. However, there was a long-term debate on the effectiveness of unlabeled data in text classification. This was partially caused by the fact that many assumptions in theoretic analysis often do not hold in practice. We argue that this problem may be further understood by adding an additional dimension in the experiment. This allows us to address this problem in the perspective of bias and variance in a broader view. We show that the well-known performance degradation issue caused by unlabeled data can be reproduced as a subset of the whole scenario. We argue that if the bias-variance tradeoff is to be better balanced by a more effective feature selection method unlabeled data is very likely to boost the classification performance. We then propose a feature selection framework in which labeled and unlabeled training samples are both considered. We discuss its potential in achieving such a balance. Besides, the application in financial sentiment analysis is chosen because it not only exemplifies an important application, the data possesses better illustrative power as well. The implications of this study in text classification and financial sentiment analysis are both discussed.

v2026.09.13