Arrow Research search

Author name cluster

Haotian Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

EAAI Journal 2026 Journal Article

Deep learning-based coke dry quenching material location prediction using physical information reconstruction features

  • Xinyang Meng
  • Keliang Pang
  • Zhiyuan Gu
  • Youzhi Zheng
  • Fujun Liu
  • Chaoran Wan
  • Haotian Wu
  • Minmin Sun

Coke dry quenching (CDQ) is a common, environmentally friendly technology applied in iron and steel production and plays an important role in improving coke quality as well as in emission reduction and pollution reduction. Material location prediction is crucial for ensuring the stable operation of dry quenching systems. In this paper, we propose a novel artificial intelligence approach for predicting the location of coke materials in CDQ furnaces by incorporating a method known as physical information feature reconstruction (PIFR). This method integrates physical a priori knowledge (such as the law of mass conservation and furnace structural characteristics) into the feature engineering process, effectively improving the accuracy and stability of time-series predictions in both single-step and multistep forecasting tasks. The experimental results demonstrate that PIFR significantly enhances the performance of various deep learning models. Specifically, for the long short-term memory model, the mean squared error and mean absolute error decreased by 51. 25% and 37. 63%, respectively, whereas the coefficient of determination increased to 0. 941. Moreover, PIFR effectively mitigates issues commonly encountered in multi-step prediction, such as cumulative error and prediction curve flattening. The application of PIFR not only improves the accuracy of the model but also significantly enhances its generalization capability.

ICLR Conference 2025 Conference Paper

Actions Speak Louder Than Words: Rate-Reward Trade-off in Markov Decision Processes

  • Haotian Wu
  • Gongpu Chen
  • Deniz Gündüz

The impact of communication on decision-making systems has been extensively studied under the assumption of dedicated communication channels. We instead consider communicating through actions, where the message is embedded into the actions of an agent which interacts with the environment in a Markov decision process (MDP) framework. We conceptualize the MDP environment as a finite-state channel (FSC), where the actions of the agent serve as the channel input, while the states of the MDP observed by another agent (i.e., receiver) serve as the channel output. Here, we treat the environment as a communication channel over which the agent communicates through its actions, while at the same time, trying to maximize its reward. We first characterize the optimal information theoretic trade-off between the average reward and the rate of reliable communication in the infinite-horizon regime. Then, we propose a novel framework to design a joint control/coding policy, termed Act2Comm, which seamlessly embeds messages into actions. From a communication perspective, Act2Comm functions as a learning-based channel coding scheme for non-differentiable FSCs under input-output constraints. From a control standpoint, Act2Comm learns an MDP policy that incorporates communication capabilities, though at the cost of some control performance. Overall, Act2Comm effectively balances the dual objectives of control and communication in this environment. Experimental results validate Act2Comm's capability to enable reliable communication while maintaining a certain level of control performance.

ICRA Conference 2025 Conference Paper

DiffCP: Ultra-Low Bit Collaborative Perception via Diffusion Model

  • Ruiqing Mao
  • Haotian Wu
  • Yukuan Jia
  • Zhaojun Nan
  • Yuxuan Sun 0001
  • Sheng Zhou 0001
  • Deniz Gündüz
  • Zhisheng Niu

Collaborative perception (CP) is emerging as a promising solution to the inherent limitations of stand-alone intelligence. However, current wireless communication systems are unable to support feature-level and raw-level collaborative algorithms due to their enormous bandwidth demands. In this paper, we propose DiffCP, a novel CP paradigm that utilizes a diffusion model to efficiently compress the sensing information of collaborators. By incorporating both geometric and semantic conditions into the generative model, DiffCP enables feature-level collaboration with an ultra-low communication cost, advancing the practical implementation of CP systems. This paradigm can be seamlessly integrated into existing CP algorithms to enhance a wide range of downstream tasks. Through extensive experimentation, we investigate the tradeoffs between communication, computation, and performance. Numerical results demonstrate that DiffCP can significantly reduce communication costs by 14. 5-fold while maintaining the same performance as the state-of-the-art algorithm.

ICML Conference 2025 Conference Paper

LotteryCodec: Searching the Implicit Representation in a Random Network for Low-Complexity Image Compression

  • Haotian Wu
  • Gongpu Chen
  • Pier Luigi Dragotti
  • Deniz Gündüz

We introduce and validate the lottery codec hypothesis, which states that untrained subnetworks within randomly initialized networks can serve as synthesis networks for overfitted image compression, achieving rate-distortion (RD) performance comparable to trained networks. This hypothesis leads to a new paradigm for image compression by encoding image statistics into the network substructure. Building on this hypothesis, we propose LotteryCodec, which overfits a binary mask to an individual image, leveraging an over-parameterized and randomly initialized network shared by the encoder and the decoder. To address over-parameterization challenges and streamline subnetwork search, we develop a rewind modulation mechanism that improves the RD performance. LotteryCodec outperforms VTM and sets a new state-of-the-art in single-image compression. LotteryCodec also enables adaptive decoding complexity through adjustable mask ratios, offering flexible compression solutions for diverse device constraints and application requirements.

NeurIPS Conference 2025 Conference Paper

MoRIC: A Modular Region-based Implicit Codec for Image Compression

  • Gen Li
  • Haotian Wu
  • Deniz Gunduz

We introduce Modular Region-Based Implicit Codec (MoRIC), a novel image compression algorithm that relies on implicit neural representations (INRs). Unlike previous INR-based codecs that model the entire image with a single neural network, MoRIC assigns dedicated models to distinct regions in the image, each tailored to its local distribution. This region-wise design enhances adaptation to local statistics and enables flexible, single-object compression with fine-grained rate-distortion (RD) control. MoRIC allows regions of arbitrary shapes, and provides the contour information for each region as separate information. In particular, it incorporates adaptive chain coding for lossy and lossless contour compression, and a shared global modulator that injects multi-scale global context into local overfitting processes in a coarse-to-fine manner. MoRIC achieves state-of-the-art performance in single-object compression with significantly lower decoding complexity than existing learned neural codecs, which results in a highly efficient compression approach for fixed-background scenarios, e. g. , for surveillance cameras. It also sets a new benchmark among overfitted codecs for standard image compression. Additionally, MoRIC naturally supports semantically meaningful layered compression through selective region refinement, paving the way for scalable and flexible INR-based codecs.

EAAI Journal 2025 Journal Article

Multisource-domain regression transfer learning framework for predicting student academic performance considering balanced similarity

  • Li Wang
  • Lucong Zhang
  • Haotian Wu
  • Teng Zhang
  • Ke Qiu
  • Tianyu Chen
  • Hongwu Qin

The increasing integration of information technology and artificial intelligence has extensively implemented computer-aided intelligent education systems in higher education. A critical task within these systems is student performance prediction, which forecasts future academic outcomes by analyzing data such as historical grades, learning behaviors, and classroom participation. This enables early intervention and personalized teaching based on scientific evidence. However, most existing methods rely on traditional machine learning techniques, which can hardly address issues such as domain distribution discrepancies and data imbalance effectively. To overcome these challenges, we propose a multisource-domain transfer learning regression framework that integrates domain selection, hybrid feature extraction, and dynamic joint distribution adaptation techniques. Specifically, the framework first selects appropriate source domains on the basis of preset thresholds via cross-validation. Thereafter, a hybrid feature extractor is used to derive (i) common features from the target and selected source domains and (ii) domain-specific features from the target domain. Finally, a dynamic adaptive factor is introduced to balance differences between the marginal and conditional distributions. Experimental results indicate that the proposed framework significantly reduces the root mean square error with an average prediction improvement of 21. 05 %, compared with baseline methods and other advanced approaches.

IROS Conference 2025 Conference Paper

Tactile sensing soft fingertip with dual air bag structure for an anthropomorphic robotic hand

  • Jipeng Yin
  • Yuchao Wang
  • Yuwei Liu
  • Haotian Wu
  • Yinian Mao
  • Qiliang Zhong
  • Ruichen Zhen
  • Yang Yang

Tactile sensing plays a crucial role to empower robotic hands with improved grasping and manipulating abilities. In this paper, we propose an anthropomorphic robotic hand design with dual air bag sensors integrated soft fingertips to achieve tactile sensing. The air bag sensor is low-cost, easy-to-build and deformable, and can be embedded in the fingertip, endows the hand with the ability to perceive and makes it have the mechanical complicance similar to the human fingertip. The air bag sensor exhibits high performance metrics, including a sensitivity of ~1. 65 kPa/N, a minimum detection force of < 0. 01 N, a response time of < 10 ms, and good stability and repeatability. The experimental results show that the proposed robotic hand performs well in surface texture detection, hard inclusion depth detection and object softness detection, as well as grasping tasks. By applying a machine learning algorithm to the experimental data, an accuracy of 0. 767 and 0. 898 was achieved in predicting hard inclusion depth and object hardness, respectively. This study provides a simple and effective tactile sensing solution for the design of anthropomorphic robotic hand, and may have possible applications such as end-effectors for humanoid robots or robotic palpation.

ICML Conference 2024 Conference Paper

Pedestrian Attribute Recognition as Label-balanced Multi-label Learning

  • Yibo Zhou
  • Hai-Miao Hu
  • Yirong Xiang
  • Xiaokang Zhang
  • Haotian Wu

Rooting in the scarcity of most attributes, realistic pedestrian attribute datasets exhibit unduly skewed data distribution, from which two types of model failures are delivered: (1) label imbalance: model predictions lean greatly towards the side of majority labels; (2) semantics imbalance: model is easily overfitted on the under-represented attributes due to their insufficient semantic diversity. To render perfect label balancing, we propose a novel framework that successfully decouples label-balanced data re-sampling from the curse of attributes co-occurrence, i. e. , we equalize the sampling prior of an attribute while not biasing that of the co-occurred others. To diversify the attributes semantics and mitigate the feature noise, we propose a Bayesian feature augmentation method to introduce true in-distribution novelty. Handling both imbalances jointly, our work achieves best accuracy on various popular benchmarks, and importantly, with minimal computational budget.

v2026.09.13