Arrow Research search

Author name cluster

Wenxin Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
1 author row

Possible papers

8

AAMAS Conference 2026 Conference Paper

HLSMAC: A New StarCraft Multi-Agent Challenge for High-Level Strategic Decision-Making

  • Xingxing Hong
  • Yungong Wang
  • Dexin Jin
  • Ye Yuan
  • Ximing Huang
  • Zijian Wu
  • Yirui Rao
  • Wenxin Li

Benchmarks are crucial for assessing multi-agent reinforcement learning (MARL) algorithms. While StarCraft II-related environments have driven significant advances in MARL, existing benchmarks like SMAC focus primarily on micromanagement, limiting comprehensiveevaluationofhigh-levelstrategicintelligence. Toaddress this, we introduce HLSMAC, a new cooperative MARL benchmark with 12 carefully designed StarCraft II scenarios based on classical stratagems from the Thirty-Six Stratagems. Each scenario corresponds to a specific stratagem and is designed to challenge agents with diverse strategic elements, including tactical maneuvering, timing coordination, and deception, thereby opening up avenues for evaluating high-level strategic decision-making capabilities. We also propose novel metrics across multiple dimensions beyond conventional win rate, such as ability utilization and advancement efficiency, to assess agents’ overall performance within the HLSMAC environment. We conduct a large-scale evaluation of 21 state-of-the-art MARL algorithms and LLM-based agents, with additional multi-seed analysis for relatively better-performing methods. The results demonstrate that HLSMAC serves as a robust testbed for advancing multi-agent strategic decision-making.

AAMAS Conference 2026 Conference Paper

Pareto-guided Pipeline for Distilling Featherweight AI Agents in Mobile MOBA Games

  • Xionghui Yang
  • Bozhou Chen
  • Yunlong Lu
  • Yongyi Wang
  • Lingfeng Li
  • Lanxiao Huang
  • Lin Liu
  • Wenjun Wang

Recent advances in game AI have demonstrated the feasibility of training agents that surpass top-tier human professionals in complex environments such as Honor of Kings (HoK), a leading mobile multiplayer online battle arena (MOBA) game. However, deploying such powerful agents on mobile devices remains a major challenge. On one hand, the intricate multi-modal state representation and hierarchical action space of HoK demand large, sophisticated policy networks that are inherently difficult to compress into lightweight forms. On the other hand, production deployment requires highfrequency inference under strict energy and latency constraints on mobile platform. To the best of our knowledge, bridging largescale game AI and practical on-device deployment has not been systematically studied. In this work, we propose a Pareto optimality guided pipeline and design a high-efficiency student architecture search space tailored for mobile execution, enabling systematic exploration of the trade-off between performance and efficiency. Experimental results demonstrate that the distilled model achieves remarkable efficiency, including an 12. 4× faster inference speed (under 0. 5ms per frame) and a 15. 6× improvement in energy efficiency (under 0. 5mAh per game), while retaining a 40. 32% win rate against the original teacher model. FullVersion: Thefull version ofthispaper, including theappendix, is available on arXiv. 1 1https: //arxiv. org/abs/2602. 07521 This work is licensed under a Creative Commons Attribution International 4. 0 License. Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), C. Amato, L. Dennis, V. Mascardi, J. Thangarajah (eds.), May 25 – 29, 2026, Paphos, Cyprus. © 2026 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org). https: //doi. org/10. 65109/HUOT2523

NeurIPS Conference 2025 Conference Paper

InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback

  • Boyuan Chen
  • Donghai Hong
  • Jiaming Ji
  • Jiacheng Zheng
  • Bowen Dong
  • Jiayi Zhou
  • Kaile Wang
  • Juntao Dai

As multimodal large models (MLLMs) continue to advance across challenging tasks, a key question emerges: \textbf{\textit{What essential capabilities are still missing? }}A critical aspect of human learning is continuous interaction with the environment -- not limited to language, but also involving multimodal understanding and generation. To move closer to human-level intelligence, models must similarly support \textbf{multi-turn}, \textbf{multimodal interaction}. In particular, they should comprehend interleaved multimodal contexts and respond coherently in ongoing exchanges. In this work, we present \textbf{an initial exploration} through the \textsc{InterMT} -- \textbf{the first preference dataset for \textit{multi-turn} multimodal interaction}, grounded in real human feedback. In this exploration, we particularly emphasize the importance of human oversight, introducing expert annotations to guide the process, motivated by the fact that current MLLMs lack such complex interactive capabilities. \textsc{InterMT} captures human preferences at both global and local levels into nine sub-dimensions, consists of 15. 6k prompts, 52. 6k multi-turn dialogue instances, and 32. 4k human-labeled preference pairs. To compensate for the lack of capability for multi-modal understanding and generation, we introduce an agentic workflow that leverages tool-augmented MLLMs to construct multi-turn QA instances. To further this goal, we introduce \textsc{InterMT-Bench} to assess the ability ofMLLMs in assisting judges with multi-turn, multimodal tasks. We demonstrate the utility of \textsc{InterMT} through applications such as judge moderation and further reveal the \textit{multi-turn scaling law} of judge model. We hope the open-source of our data can help facilitate further research on aligning current MLLMs to the next step.

EAAI Journal 2024 Journal Article

A user review data-driven supplier ranking model using aspect-based sentiment analysis and fuzzy theory

  • Bingli Sun
  • Xiao Song
  • Wenxin Li
  • Lu Liu
  • Guanghong Gong
  • Yan Zhao

Background: The supplier selection problem is a sophisticated decision-making process that involves evaluating multiple factors. While previous research has primarily focused on objective attributes, such as supplier qualifications, product quality, and price, the subjective opinions of users have often been overlooked. However, with the growing importance of user reviews and sentiment analysis in e-commerce, incorporating users’ opinions on supplier products can provide valuable insights. Purpose: This study aims to address the limitations of existing supplier selection approaches by proposing a comprehensive framework that integrates aspect-level sentiment analysis and a fuzzy multi-attribute decision model. The goal is to enhance the decision-making process by considering both objective attributes and subjective opinions. Methods: To achieve this, we develop a novel convolutional neural network (CNN) model with a gating mechanism to perform aspect-level sentiment analysis. Furthermore, we propose a fuzzy multi-attribute decision model that combines the predefined sentiment aspects with traditional evaluation criteria. The model is applied to a dataset specifically designed for automotive component supplier selection. Results: Experimental results demonstrate the superior performance of our approach compared to existing methods and datasets. A case study demonstrates the combination of aspect-level sentiment analysis and the fuzzy decision model allows for a more comprehensive evaluation of suppliers. Conclusion: By integrating aspect-level sentiment analysis and the fuzzy multi-attribute decision model, our proposed framework offers a novel perspective on supplier selection problems. The results highlight the feasibility and superiority of our approach, providing valuable insights for management in making informed decisions. This research contributes to the fields of supplier selection, sentiment analysis, and decision-making, with potential applications in various industries beyond the automotive sector.

IJCAI Conference 2024 Conference Paper

Mahjong AI Competition: Exploring AI Application in Complex Real-World Games

  • Yunlong Lu
  • Wenxin Li

This paper presents three Mahjong AI competitions we held at IJCAI. We briefly introduce the rule of Mahjong and its challenges to AI algorithms. By showing the results and the application of various algorithms in the competitions, we claim that existing algorithms show promising results in Mahjong, while open problems remain and more efforts are needed towards solving this complex game.

NeurIPS Conference 2022 Conference Paper

Submodular Maximization in Clean Linear Time

  • Wenxin Li
  • Moran Feldman
  • Ehsan Kazemi
  • Amin Karbasi

In this paper, we provide the first deterministic algorithm that achieves $1/2$-approximation for monotone submodular maximization subject to a knapsack constraint, while making a number of queries that scales only linearly with the size of the ground set $n$. Moreover, our result automatically paves the way for developing a linear-time deterministic algorithm that achieves the tight $1-1/e$ approximation guarantee for monotone submodular maximization under a cardinality (size) constraint. To complement our positive results, we also show strong information-theoretic lower bounds. More specifically, we show that when the maximum cardinality allowed for a solution is constant, no deterministic or randomized algorithm making a sub-linear number of function evaluations can guarantee any constant approximation ratio. Furthermore, when the constraint allows the selection of a constant fraction of the ground set, we show that any algorithm making fewer than $\Omega(n/\log(n))$ function evaluations cannot perform better than an algorithm that simply outputs a uniformly random subset of the ground set of the right size. We extend our results to the general case of maximizing a monotone submodular function subject to the intersection of a $p$-set system and multiple knapsack constraints. Finally, we evaluate the performance of our algorithms on multiple real-life applications, including movie recommendation, location summarization, Twitter text summarization, and video summarization.

IJCAI Conference 2019 Conference Paper

Optimizing Constraint Solving via Dynamic Programming

  • Shu Lin
  • Na Meng
  • Wenxin Li

Constraint optimization problems (COP) on finite domains are typically solved via search. Many problems (e. g. , 0-1 knapsack) involve redundant search, making a general constraint solver revisit the same subproblems again and again. Existing approaches use caching, symmetry breaking, subproblem dominance, or search with decomposition to prune the search space of constraint problems. In this paper we present a different approach--DPSolver--which uses dynamic programming (DP) to efficiently solve certain types of constraint optimization problems (COPs). Given a COP modeled with MiniZinc, DPSolver first analyzes the model to decide whether the problem is efficiently solvable with DP. If so, DPSolver refactors the constraints and objective functions to model the problem as a DP problem. Finally, DPSolver feeds the refactored model to Gecode--a widely used constraint solver--for the optimal solution. Our evaluation shows that DPSolver significantly improves the performance of constraint solving.

IJCAI Conference 2018 Conference Paper

Learning to Design Games: Strategic Environments in Reinforcement Learning

  • Haifeng Zhang
  • Jun Wang
  • Zhiming Zhou
  • Weinan Zhang
  • Yin Wen
  • Yong Yu
  • Wenxin Li

In typical reinforcement learning (RL), the environment is assumed given and the goal of the learning is to identify an optimal policy for the agent taking actions through its interactions with the environment. In this paper, we extend this setting by considering the environment is not given, but controllable and learnable through its interaction with the agent at the same time. This extension is motivated by environment design scenarios in the real-world, including game design, shopping space design and traffic signal design. Theoretically, we find a dual Markov decision process (MDP) w. r. t. the environment to that w. r. t. the agent, and derive a policy gradient solution to optimizing the parametrized environment. Furthermore, discontinuous environments are addressed by a proposed general generative framework. Our experiments on a Maze game design task show the effectiveness of the proposed algorithms in generating diverse and challenging Mazes against various agent settings.

v2026.09.13