Arrow Research search

Author name cluster

Jin Huang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

30 papers
2 author rows

Possible papers

30

AAAI Conference 2026 Conference Paper

A Better Start: Sensitivity-Aware Warm-Up for Robust and Efficient Fine-Tuning

  • Yile Chen
  • Zeyi Wen
  • Jian Chen
  • Jin Huang

As an essential component of fine-tuning, warm-up plays a crucial role in promoting stability and generalization. Many studies have examined its underlying mechanisms from different aspects. However, most of the studies focus on incorporating these insights into optimizers to reduce the reliance on warm-up. Little attention has been paid to addressing the inherent limitations of the warm-up itself, which restricts its effectiveness. In this work, we revisit warm-up from a loss landscape perspective and identify several limitations with existing warm-up, including: (1) susceptibility to nearby suboptimal traps, (2) sensitivity to hyperparameters and random seeds, and (3) inefficiency during the early stages of training. To overcome these limitations, we propose Sensitivity-Aware Warm-Up (SAWU), a lightweight and adaptive strategy that dynamically leverages learning sensitivity during warm-up to guide updates toward better and more stable basins. In addition, SAWU also introduces an adaptive scheduling mechanism and phase transition strategy across warm-up, stable, and decay phases to further enhance robustness and efficiency. Extensive experiments on various downstream tasks show that SAWU significantly outperforms the vanilla method (e.g., average 3.43% improvement on RoBerta). Moreover, SAWU can be easily combined with various optimizers and remains effective even when warm-up-based methods fail (e.g, it lifts RAdam from 49.46% to 91.78% on qnli. Thanks to its lightweight nature, SAWU introduces minimal overhead and even reduces training time by over 5% compared to other methods.

AAMAS Conference 2026 Conference Paper

D^3MAS: Decompose, Deduce, and Distribute for Enhanced Knowledge Sharing in Multi-Agent Systems

  • Heng Zhang
  • Yuling Shi
  • Xiaodong Gu
  • Haochen You
  • Zijian Zhang
  • Lubin Gan
  • Yilei Yuan
  • Jin Huang

Multi-agent systems powered by large language models exhibit strongcapabilitiesincollaborativeproblem-solving. However, these systems suffer from substantial knowledge redundancy. Agents duplicate efforts in retrieval and reasoning processes. This inefficiency stems from a deeper issue: current architectures lack mechanisms to ensure agents share minimal sufficient information at each operational stage. Empirical analysis reveals an average knowledge duplication rate of 47. 3% across agent communications. We propose D3MAS (Decompose, Deduce, and Distribute), a hierarchical coordination framework addressing redundancy through structural design rather than explicit optimization. The framework organizes collaboration across three coordinated layers. Task decomposition filters irrelevant sub-problems early. Collaborative reasoning captures complementary inference paths across agents. Distributed memoryprovidesaccesstonon-redundantknowledge. Theselayers coordinate through structured message passing in a unified heterogeneous graph. This cross-layer alignment ensures information remains aligned with actual task needs. Experiments on four challenging datasets show that D3MAS consistently improves reasoning accuracy by 8. 7% to 15. 6% and reduces knowledge redundancy by 46% on average.

AAMAS Conference 2026 Conference Paper

HyperAgent: Leveraging Hypergraphs for Topology Optimization in Multi-Agent Communication

  • Heng Zhang
  • Yuling Shi
  • Xiaodong Gu
  • Zijian Zhang
  • Haochen You
  • Lubin Gan
  • Yilei Yuan
  • Jin Huang

Recent advances in large language model-powered multi-agent systems have demonstrated remarkable collective intelligence through effective communication. However, existing approaches face two primary challenges: (i) Ineffective group collaboration modeling, as they rely on pairwise edge representations in graph structures, limiting their ability to capture relationships among multiple agents; and (ii) Limited task-adaptiveness in communication topology design, leading to excessive communication cost for simple tasks and insufficient coordination for complex scenarios. These issues restrict the scalability and practical deployment of adaptive collaboration frameworks. To address these challenges, we propose HyperAgent, a hypergraph-based framework that optimizes communication topologies and effectively captures group collaboration patterns using direct hyperedge representations. Unlike edge-based approaches, HyperAgent uses hyperedges to link multiple agents within the same subtask and employs hypergraph convolutional layers to achieve one-step information aggregation in collaboration groups. Additionally, it incorporates a variational autoencoder framework with sparsity regularization to dynamically adjust hypergraph topologies based on task complexity. Experiments highlight the superiority of HyperAgent in both performance and efficiency. For instance, on GSM8K, HyperAgent achieves 95. 07% accuracy while reducing token consumption by 25. 33%, demonstrating the potential of hypergraph-based optimization for multi-agent communication. ∗Corresponding author. This work is licensed under a Creative Commons Attribution International 4. 0 License. Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), C. Amato, L. Dennis, V. Mascardi, J. Thangarajah (eds.), May 25 – 29, 2026, Paphos, Cyprus. © 2026 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org). https: //doi. org/10. 65109/QTVF9552

JBHI Journal 2026 Journal Article

Rethinking Multi-center Semi-supervised Breast Cancer Ultrasound Image Segmentation: An Intermediate-domain Perspective

  • Zhaoyi Ye
  • Yimin Zhang
  • Jin Huang
  • Du Wang
  • Sheng Liu
  • Liye Mei
  • Cheng Lei

Multi-center breast ultrasound images eg mentation aims to leverage limited labeled data from a single center to enhance model discriminability across unlabeled data from other centers. However, differences in equipment parameters, disease severity, and imaging conditions collectively contribute to significant cross domain shifts in multi-center data. In a spirit of the golden mean, we argue that constructing an intermediate domain between the source and target domains can effectively improve model generalization. Therefore, we propose a Cross-domain Few-label Generalization (CFG) framework for multi-center breast ultrasound image segmentation. Specifically, we design the Intermediate Domain Generator (IDG) to generate intermediate domain samplesthatcontain features from both the source and target domains bidirectionally, enabling the model to explicitly learn univer sal semantic representations. Additionally, we apply Swin Masked Autoencoder (MAE) to mask and reconstruct ul trasound images, simulating speckle noise encountered during clinical ultrasound acquisition, thereby increasing the diversity of intermediate domain samples. Further more, we integrate the Kolmogorov-Arnold Network (KAN) with UNet to construct KAN-UNet, integrating learnable spline functions directly onto the edges, enabling effective multi-scale perception of breast cancer lesion features. Experimental results show that even with limited labeled data from the source domain (BUSI-WHU), the CFG frame work achieves a Kappa value of 77. 17%, surpassing ten state-of-the-art methods and outperforming the second best method by 0. 78% across four multi-center ultrasound datasets (BUSI-WHU, BUSI, Dataset-B, and Dataset-C) collected from different medical centers. The code is available at https://github.com/yzygit1230/CFG.

EAAI Journal 2026 Journal Article

Sound zoning modulation within non-enclosed vehicular cavities: targeted compensation for dynamic reverberation and wind-induced noise

  • Xu Li
  • Xudong Wu
  • Jin Huang
  • Qitao Feng

Sound zoning facilitates the creation of multiple isolated listening regions within shared physical spaces, demonstrating considerable potential for personalized audio delivery. Nevertheless, practical implementation in vehicular cavities encounters distinctive challenges characterized by dynamically changing reverberant environments and accompanying wind-induced noise during window operation. To address this limitation, this study proposes a modulation methodology of sound zoning for non-enclosed vehicular cavities under varying window openings, with a focus on characterizing sound propagation with dynamic reverberation and wind-induced noise. A physics-informed neural network framework adapted to dynamic reverberant environments enables efficient and accurate prediction of acoustic transfer functions (ATFs) under varying window operations, circumventing the high cost of repeated simulations or measurements. Furthermore, by implementing tailored assignment of expected amplitude in acoustically shielded region, supported by established ATFs prediction, a targeted compensation strategy for sound zoning modulation adapted to non-enclosed vehicular cavities is developed. Eventually, to evaluate the effectiveness of sound zoning modulation in non-enclosed cavities, systematic validation is performed through reverberant sound propagation analysis, experimental ATFs measurements, and wind tunnel testing. Results demonstrate a root mean square error below 2 dB and prediction errors under 1 dB in low-frequency ranges, while maintaining consistent full-frequency performance. More significantly, the proposed compensation not only maintains acoustic privacy between regions but also reduces computational and measuring costs.

JAIR Journal 2026 Journal Article

Understanding AI Trustworthiness: A Scoping Review of AIES & FAccT Articles

  • Siddharth Mehrotra
  • Jin Huang
  • Xuelong Fu
  • Roel Dobbe
  • Clara I. Sánchez
  • Maarten de Rijke

Background: Trustworthy AI serves as a foundational pillar for two major AI ethics conferences: AIES and FAccT. Current research often adopts techno-centric approaches, focusing primarily on technical attributes such as accuracy, reliability, robustness, and fairness, while overlooking the sociotechnical dimensions critical to understanding AI trustworthiness in real-world contexts. Objectives: This scoping review aims to examine how the AIES and FAccT communities conceptualize, measure, and validate AI trustworthiness, identifying major gaps and opportunities for advancing a holistic understanding of trustworthy AI systems. Methods: We conduct a scoping review of the AIES and FAccT conference proceedings to date, systematically analyzing how trustworthiness is defined, operationalized, and applied across different research domains. Our analysis focuses on conceptualization approaches, measurement methods, verification and validation techniques, application areas, and underlying values. Results: While significant progress has been made in defining technical attributes such as transparency, accountability, and robustness, our findings reveal critical gaps. Current research often predominantly emphasizes technical precision at the expense of social and ethical considerations. The sociotechnical nature of AI systems remains less explored and trustworthiness emerges as a contested concept shaped by those with the power to define it. Conclusions: An interdisciplinary approach combining technical rigor with social, cultural, and institutional considerations is essential for advancing trustworthy AI. We propose actionable measures for the AI ethics community to adopt holistic frameworks that genuinely address the complex interplay between AI systems and society, ultimately promoting responsible technological development that benefits all stakeholders.

IROS Conference 2025 Conference Paper

Automatic MILP Model Construction for Multi-Robot Task Allocation and Scheduling Based on Large Language Models

  • Mingming Peng
  • Zhendong Chen
  • Jie Yang
  • Jin Huang
  • Zhengqi Shi
  • Qihao Liu
  • Xinyu Li 0001
  • Liang Gao 0001

With the accelerated development of Industry 4. 0, intelligent manufacturing systems increasingly require efficient task allocation and scheduling in multi-robot systems. However, existing methods rely on domain expertise and face challenges in adapting to dynamic production constraints. Additionally, enterprises have high privacy requirements for production scheduling data, which prevents the use of cloud-based large language models (LLMs) for solution development. To address these challenges, there is an urgent need for an automated modeling solution that meets data privacy requirements. This study proposes a knowledge-augmented mixed integer linear programming (MILP) automated formulation framework, integrating local LLMs with domain-specific knowledge bases to generate executable code from natural language descriptions automatically. The framework employs a knowledge-guided DeepSeek-R1-Distill-Qwen-32B model to extract complex spatiotemporal constraints (82% average accuracy) and leverages a supervised fine-tuned Qwen2. 5-Coder-7B-Instruct model for efficient MILP code generation (90% average accuracy). Experimental results demonstrate that the framework successfully achieves automatic modeling in the aircraft skin manufacturing case while ensuring data privacy and computational efficiency. This research provides a low-barrier and highly reliable technical path for modeling in complex industrial scenarios.

JBHI Journal 2025 Journal Article

EMGANet: Edge-Aware Multi-Scale Group-Mix Attention Network for Breast Cancer Ultrasound Image Segmentation

  • Jin Huang
  • Yazhao Mao
  • Jingwen Deng
  • Zhaoyi Ye
  • Yimin Zhang
  • Jingwen Zhang
  • Lan Dong
  • Hui Shen

Breast cancer is one of the most prevalent diseases for women worldwide. Early and accurate ultrasound image segmentation plays a crucial role in reducing mortality. Although deep learning methods have demonstrated remarkable segmentation potential, they still struggle with challenges in ultrasound images, including blurred boundaries and speckle noise. To generate accurate ultrasound image segmentation, this paper proposes the Edge-Aware Multi-Scale Group-Mix Attention Network (EMGANet), which generates accurate segmentation by integrating deep and edge features. The Multi-Scale Group Mix Attention block effectively aggregates both sparse global and local features, ensuring the extraction of valuable information. The subsequent Edge Feature Enhancement block then focuses on cancer boundaries, enhancing the segmentation accuracy. Therefore, EMGANet effectively tackles unclear boundaries and noise in ultrasound images. We conduct experiments on two public datasets (Dataset-B, BUSI) and one private dataset which contains 927 samples from Renmin Hospital of Wuhan University (BUSI-WHU). EMGANet demonstrates superior segmentation performance, achieving an overall accuracy (OA) of 98. 56%, a mean IoU (mIoU) of 90. 32%, and an ASSD of 6. 1 pixels on the BUSI-WHU dataset. Additionally, EMGANet performs well on two public datasets, with a mIoU of 88. 2% and an ASSD of 9. 2 pixels on Dataset-B, and a mIoU of 81. 37% and an ASSD of 18. 27 pixels on the BUSI dataset. EMGANet achieves a state-of-the-art segmentation performance of about 2% in mIoU across three datasets. In summary, the proposed EMGANet significantly improves breast cancer segmentation through Edge-Aware and Group-Mix Attention mechanisms, showing great potential for clinical applications.

IJCAI Conference 2025 Conference Paper

Human Activity Recognition in an Open World (Abstract Reprint)

  • Derek Prijatelj
  • Samuel Grieggs
  • Jin Huang
  • Dawei Du
  • Ameya Shringi
  • Christopher Funk
  • Adam Kaufman
  • Eric Robertson

Managing novelty in perception-based human activity recognition (HAR) is critical in realistic settings to improve task performance over time and ensure solution generalization outside of prior seen samples. Novelty manifests in HAR as unseen samples, activities, objects, environments, and sensor changes, among other ways. Novelty may be task-relevant, such as a new class or new features, or task-irrelevant resulting in nuisance novelty, such as never before seen noise, blur, or distorted video recordings. To perform HAR optimally, algorithmic solutions must be tolerant to nuisance novelty, and learn over time in the face of novelty. This paper 1) formalizes the definition of novelty in HAR building upon the prior definition of novelty in classification tasks, 2) proposes an incremental open world learning (OWL) protocol and applies it to the Kinetics datasets to generate a new benchmark KOWL-718, 3) analyzes the performance of current stateof-the-art HAR models when novelty is introduced over time, 4) provides a containerized and packaged pipeline for reproducing the OWL protocol and for modifying for any future updates to Kinetics. The experimental analysis includes an ablation study of how the different models perform under various conditions as annotated by Kinetics-AVA. The code may be used to analyze different annotations and subsets of the Kinetics datasets in an incremental open world fashion, as well as be extended as further updates to Kinetics are released.

JBHI Journal 2025 Journal Article

MRRM: Advanced Biomarker Alignment in Multi-Staining Pathology Images via Multi-Scale Ring Rotation-Invariant Matching

  • Xiaoxiao Li
  • Taobo Hu
  • Zhengxiong Li
  • Mengping Long
  • Zhaoyi Ye
  • Jin Huang
  • Yaxiaer Yalikun
  • Sheng Liu

Pathology image matching is crucial for assisting pathologists in the comprehensive diagnosis of cancerous areas. However, variations in image rotation and staining caused by inherent slide imaging techniques increase the burden on pathologists, complicating the examination of cancer across different pathology slides. To address this challenge, we introduce multi-scale ring rotation-invariant matching (MRRM), which improves image matching efficiency using ring topology, assisting pathologists in robustly aligning biomarker information across various pathology images. Specifically, by employing multi-scale rings as convolution kernels, we accurately locate keypoints from the differencing of the ring pyramid, which not only enhances the likelihood of successful pathology image matching but also supports our feature descriptor in achieving advantageous performance in rotation-invariance. Experiments show that with manually annotated golden landmarks as the standard in 81 cases, exhibiting significantly superior matching accuracy (130. 93 $\, \mu \mathrm{m}$ ) and a success rate of 93. 83% compared to other methods, particularly in cases with rotated pathology images. This meets the routine diagnostic requirements of pathologists for cancer diagnosis.

TMLR Journal 2024 Journal Article

Can LLMs Effectively Leverage Graph Structural Information through Prompts, and Why?

  • Jin Huang
  • Xingjian Zhang
  • Qiaozhu Mei
  • Jiaqi Ma

Large language models (LLMs) are gaining increasing attention for their capability to process graphs with rich text attributes, especially in a zero-shot fashion. Recent studies demonstrate that LLMs obtain decent text classification performance on common text-rich graph benchmarks, and the performance can be improved by appending encoded structural information as natural languages into prompts. We aim to understand why the incorporation of structural information inherent in graph data can improve the prediction performance of LLMs. First, we rule out the concern of data leakage by curating a novel leakage-free dataset and conducting a comparative analysis alongside a previously widely-used dataset. Second, as past work usually encodes the ego-graph by describing the graph structure in natural language, we ask the question: do LLMs understand the prompts in graph structures? Third, we investigate why LLMs can improve their performance after incorporating structural information. Our exploration of these questions reveals that (i) there is no substantial evidence that the performance of LLMs is significantly attributed to data leakage; (ii) instead of understanding prompts as graph structures, LLMs tend to process prompts more as contextual paragraphs and (iii) the most efficient elements of the local neighborhood included in the prompt are phrases that are pertinent to the node label, rather than the graph structure.

AIIM Journal 2024 Journal Article

Healthcare facilities management: A novel data-driven model for predictive maintenance of computed tomography equipment

  • Haopeng Zhou
  • Qilin Liu
  • Haowen Liu
  • Zhu Chen
  • Zhenlin Li
  • Yixuan Zhuo
  • Kang Li
  • Changxi Wang

Background The breakdown of healthcare facilities is a huge challenge for hospitals. Medical images obtained by Computed Tomography (CT) provide information about the patients' physical conditions and play a critical role in diagnosis of disease. To deliver high-quality medical images on time, it is essential to minimize the occurrence frequencies of anomalies and failures of the equipment. Methods We extracted the real-time CT equipment status time series data such as oil temperature, of three equipment, between May 19, 2020, and May 19, 2021. Tube arcing is treated as the classification label. We propose a dictionary-based data-driven model SAX-HCBOP, where the two methods, Histogram-based Information Gain Binning (HIGB) and Coefficient improved Bag of Pattern (CoBOP), are implemented to transform the data into the bag-of-words paradigm. We compare our model to the existing predictive maintenance models based on statistical and time series classification algorithms. Results The results show that the Accuracy, Recall, Precision and F1-score of the proposed model achieve 0. 904, 0. 747, 0. 417, 0. 535, respectively. The oil temperature is identified as the most important feature. The proposed model is superior to other models in predicting CT equipment anomalies. In addition, experiments on the public dataset also demonstrate the effectiveness of the proposed model. Conclusions The two proposed methods can improve the performance of the dictionary-based time series classification methods in predictive maintenance. In addition, based on the proposed real-time anomaly prediction system, the model assists hospitals in making accurate healthcare facilities maintenance decisions.

JAIR Journal 2024 Journal Article

Human Activity Recognition in an Open World

  • Derek S. Prijatelj
  • Samuel Grieggs
  • Jin Huang
  • Dawei Du
  • Ameya Shringi
  • Christopher Funk
  • Adam Kaufman
  • Eric Robertson

Managing novelty in perception-based human activity recognition (HAR) is critical in realistic settings to improve task performance over time and ensure solution generalization outside of prior seen samples. Novelty manifests in HAR as unseen samples, activities, objects, environments, and sensor changes, among other ways. Novelty may be task-relevant, such as a new class or new features, or task-irrelevant resulting in nuisance novelty, such as never before seen noise, blur, or distorted video recordings. To perform HAR optimally, algorithmic solutions must be tolerant to nuisance novelty, and learn over time in the face of novelty. This paper 1) formalizes the definition of novelty in HAR building upon the prior definition of novelty in classification tasks, 2) proposes an incremental open world learning (OWL) protocol and applies it to the Kinetics datasets to generate a new benchmark KOWL-718, 3) analyzes the performance of current stateof-the-art HAR models when novelty is introduced over time, 4) provides a containerized and packaged pipeline for reproducing the OWL protocol and for modifying for any future updates to Kinetics. The experimental analysis includes an ablation study of how the different models perform under various conditions as annotated by Kinetics-AVA. The code may be used to analyze different annotations and subsets of the Kinetics datasets in an incremental open world fashion, as well as be extended as further updates to Kinetics are released.

JBHI Journal 2024 Journal Article

MSGM: An Advanced Deep Multi-Size Guiding Matching Network for Whole Slide Histopathology Images Addressing Staining Variation and Low Visibility Challenges

  • Xiaoxiao Li
  • Zhengxiong Li
  • Taobo Hu
  • Mengping Long
  • Xiao Ma
  • Jin Huang
  • Yiqiang Liu
  • Yaxiaer Yalikun

Matching whole slide histopathology images to provide comprehensive information on homologous tissues is beneficial for cancer diagnosis. However, the challenge arises with the Giga-pixel whole slide images (WSIs) when aiming for high-accuracy matching. Learning-based methods are difficult to generalize well with large-size WSIs, necessitating the integration of traditional matching methods to enhance accuracy as the size increases. In this paper, we propose a multi-size guiding matching method applicable high-accuracy requirements. Specifically, we design learning multiscale texture to train deep descriptors, called TDescNet, that trains 64 × 64 × 256 and 256 × 256 × 128 size convolution layer as C64 and C256 descriptors to overcome staining variation and low visibility challenges. Furthermore, we develop the 3D-ring descriptor using sparse keypoints to support the description of large-size WSIs. Finally, we employ C64, C256, and 3D-ring descriptors to progressively guide refined local matching, utilizing geometric consistency to identify correct matching results. Experiments show that when matching WSIs of size 4096 × 4096 pixels, our average matching error is 123. 48 μm and the success rate is 93. 02 $\%$ in 43 cases. Notably, our method achieves an average improvement of 65. 52 μm in matching accuracy compared to recent state-of-the-art methods, with enhancements ranging from 36. 27 μm to 131. 66 μm. Therefore, we achieve high-fidelity whole-slice image matching, and overcome staining variation and low visibility challenges, enabling assistance in comprehensive cancer diagnosis through matched WSIs.

AAAI Conference 2024 Conference Paper

PMRC: Prompt-Based Machine Reading Comprehension for Few-Shot Named Entity Recognition

  • Jin Huang
  • Danfeng Yan
  • Yuanqiang Cai

The prompt-based method has been proven effective in improving the performance of pre-trained language models (PLMs) on sentence-level few-shot tasks. However, when applying prompting to token-level tasks such as Named Entity Recognition (NER), specific templates need to be designed, and all possible segments of the input text need to be enumerated. These methods have high computational complexity in both training and inference processes, making them difficult to apply in real-world scenarios. To address these issues, we redefine the NER task as a Machine Reading Comprehension (MRC) task and incorporate prompting into the MRC framework. Specifically, we sequentially insert boundary markers for various entity types into the templates and use these markers as anchors during the inference process to differentiate entity types. In contrast to the traditional multi-turn question-answering extraction in the MRC framework, our method can extract all spans of entity types in one round. Furthermore, we propose word-based template and example-based template that enhance the MRC framework's perception of entity start and end positions while significantly reducing the manual effort required for template design. It is worth noting that in cross-domain scenarios, PMRC does not require redesigning the model architecture and can continue training by simply replacing the templates to recognize entity types in the target domain. Experimental results demonstrate that our approach outperforms state-of-the-art models in low-resource settings, achieving an average performance improvement of +5.2% in settings where access to source domain data is limited. Particularly, on the ATIS dataset with a large number of entity types and 10-shot setting, PMRC achieves a performance improvement of +15.7%. Moreover, our method achieves a decoding speed 40.56 times faster than the template-based cloze-style approach.

JAIR Journal 2023 Journal Article

Decentralized Gradient-Quantization Based Matrix Factorization for Fast Privacy-Preserving Point-of-Interest Recommendation

  • Xuebin Zhou
  • Zhibin Hu
  • Jin Huang
  • Jian Chen

With the rapidly growing of location-based social networks, point-of-interest (POI) recommendation has been attracting tremendous attentions. Previous works for POI recommendation usually use matrix factorization (MF)-based methods, which achieve promising performance. However, existing MF-based methods suffer from two critical limitations: (1) Privacy issues: all users’ sensitive data are collected to the centralized server which may leak on either the server side or during transmission. (2) Poor resource utilization and training efficiency: training on centralized server with potentially huge low-rank matrices is computational inefficient. In this paper, we propose a novel decentralized gradient-quantization based matrix factorization (DGMF) framework to address the above limitations in POI recommendation. Compared with the centralized MF methods which store all sensitive data and low-rank matrices during model training, DGMF treats each user’s device (e.g., phone) as an independent learner and keeps the sensitive data on each user’s end. Furthermore, a privacy-preserving and communication-efficient mechanism with gradient-quantization technique is presented to train the proposed model, which aims to handle the privacy problem and reduces the communication cost in the decentralized setting. Theoretical guarantees of the proposed algorithm and experimental studies on real-world datasets demonstrate the effectiveness of the proposed algorithm.

EAAI Journal 2023 Journal Article

Discrete limited attentional collaborative filtering for fast social recommendation

  • Zhibin Hu
  • Xuebin Zhou
  • Zhiwei He
  • Zehang Yang
  • Jian Chen
  • Jin Huang

Over the last few years, social recommendation has attracted tremendous attention due to the ever-growing online social platform such as Twitter and Facebook. However, as the number of users increases rapidly, recommendation efficiency has become the bottleneck of many existing social recommender systems due to the computation and storage of real-valued models. For addressing the efficiency problem, recent researches resolve it by introducing hashing technique into social recommender systems. By mapping real values to discrete values, the computational speed is guaranteed as well as the storage cost is reduced. Nevertheless, these methods suffer from two critical limitations: (1) The inevitable quantization loss brought by hash function decreases recommendation accuracy to a certain extent. (2) The original social relations contain massive noise that may result in sub-optimal accuracy of recommendation without considering the fact that people can only pay attention to a small number of their friends. Therefore, to tackle the above limitations and have a better tradeoff between accuracy and efficiency, in this paper, we propose a novel social recommendation method called Discrete Limited Attentional Collaborative Filtering (DLACF), which models recommendation objective with limited attention as a constrained mix-integer optimization problem. Since the original problem is NP-hard, we further devise a computationally efficient optimization algorithm to learn the binary codes as well as to estimate the best influential friends. Experimental results conducted on two real-world datasets demonstrate the effectiveness of our proposed model, achieving the averaged improvement of 118. 7% and 54. 7% compared to state-of-the-art discrete methods.

ICML Conference 2023 Conference Paper

HarsanyiNet: Computing Accurate Shapley Values in a Single Forward Propagation

  • Lu Chen
  • Siyu Lou
  • Keyan Zhang
  • Jin Huang
  • Quanshi Zhang

The Shapley value is widely regarded as a trustworthy attribution metric. However, when people use Shapley values to explain the attribution of input variables of a deep neural network (DNN), it usually requires a very high computational cost to approximate relatively accurate Shapley values in real-world applications. Therefore, we propose a novel network architecture, the HarsanyiNet, which makes inferences on the input sample and simultaneously computes the exact Shapley values of the input variables in a single forward propagation. The HarsanyiNet is designed on the theoretical foundation that the Shapley value can be reformulated as the redistribution of Harsanyi interactions encoded by the network.

AAAI Conference 2018 Conference Paper

Energy-Efficient Automatic Train Driving by Learning Driving Patterns

  • Jin Huang
  • Yue Gao
  • Sha Lu
  • Xibin Zhao
  • Yangdong Deng
  • Ming Gu

Railway is regarded as the most sustainable means of modern transportation. With the fast-growing of fleet size and the railway mileage, the energy consumption of trains is becoming a serious concern globally. The nature of railway offers a unique opportunity to optimize the energy efficiency of locomotives by taking advantage of the undulating terrains along a route. The derivation of an energy-optimal train driving solution, however, proves to be a significant challenge due to the high dimension, nonlinearity, complex constraints, and timevarying characteristic of the problem. An optimized solution can only be attained by considering both the complex environmental conditions of a given route and the inherent characteristics of a locomotive. To tackle the problem, this paper employs a high-order correlation learning method for online generation of the energy optimized train driving solutions. Based on the driving data of experienced human drivers, a hypergraph model is used to learn the optimal embedding from the specified features for the decision of a driving operation. First, we design a feature set capturing the driving status. Next all the training data are formulated as a hypergraph and an inductive learning process is conducted to obtain the embedding matrix. The hypergraph model can be used for real-time generation of driving operation. We also proposed a reinforcement updating scheme, which offers the capability of sustainable enhancement on the hypergraph model in industrial applications. The learned model can be used to determine an optimized driving operation in real-time tested on the Hardware-in-Loop platform. Validation experiments proved that the energy consumption of the proposed solution is around 10% lower than that of average human drivers.

AAAI Conference 2018 Conference Paper

Hypergraph Learning With Cost Interval Optimization

  • Xibin Zhao
  • Nan Wang
  • Heyuan Shi
  • Hai Wan
  • Jin Huang
  • Yue Gao

In many classification tasks, the misclassification costs of different categories usually vary significantly. Under such circumstances, it is essential to identify the importance of different categories and thus assign different misclassification losses in many applications, such as medical diagnosis, saliency detection and software defect prediction. However, we note that it is infeasible to determine the accurate cost value without great domain knowledge. In most common cases, we may just have the information that which category is more important than the other categories, i. e. , the identification of defect-prone softwares is more important than that of defect-free. To tackle these issues, in this paper, we propose a hypergraph learning method with cost interval optimization, which is able to handle cost interval when data is formulated using the high-order relationships. In this way, data correlations are modeled by a hypergraph structure, which has the merit to exploit the underlying relationships behind the data. With a cost-sensitive hypergraph structure, in order to improve the performance of the classifier without precise cost value, we further introduce cost interval optimization to hypergraph learning. In this process, the optimization on cost interval achieves better performance instead of choosing uncertain fixed cost in the learning process. To evaluate the effectiveness of the proposed method, we have conducted experiments on two groups of dataset, i. e. , the NASA Metrics Data Program (NASA) dataset and UCI Machine Learning Repository (UCI) dataset. Experimental results and comparisons with state-of-the-art methods have exhibited better performance of our proposed method.

AAAI Conference 2018 Short Paper

Selecting Proper Multi-Class SVM Training Methods

  • Yawen Chen
  • Zeyi Wen
  • Jian Chen
  • Jin Huang

Support Vector Machines (SVMs) are excellent candidate solutions to solving multi-class problems, and multi-class SVMs can be trained by several different methods. Different training methods commonly produce SVMs with different effectiveness, and no multi-class SVM training method always outperforms other multi-class SVM training methods on all problems. This raises difficulty for practitioners to choose the best training method for a given problem. In this work, we propose a Multi-class Method Selection (MMS) approach to help users select the most appropriate method among one-versus-one (OVO), one-versus-all (OVA) and structural SVMs (SSVMs) for a given problem. Our key idea is to select the training method based on the distribution of training data and the similarity between different classes. Using the distribution and class similarity, we estimate the unclassifiable rate of each multi-class SVM training method, and select the training method with the minimum unclassifiable rate. Our initial findings show: (i) SSVMs with linear kernel perform worse than OVO and OVA; (ii) MMS often produces SVM classifiers that can confidently classify unseen instances.

AAAI Conference 2015 Conference Paper

On Machine Learning towards Predictive Sales Pipeline Analytics

  • Junchi Yan
  • Chao Zhang
  • Hongyuan Zha
  • Min Gong
  • Changhua Sun
  • Jin Huang
  • Stephen Chu
  • Xiaokang Yang

Sales pipeline win-propensity prediction is fundamental to effective sales management. In contrast to using subjective human rating, we propose a modern machine learning paradigm to estimate the winpropensity of sales leads over time. A profile-specific two-dimensional Hawkes processes model is developed to capture the influence from seller’s activities on their leads to the win outcome, coupled with lead’s personalized profiles. It is motivated by two observations: i) sellers tend to frequently focus their selling activities and efforts on a few leads during a relatively short time. This is evidenced and reflected by their concentrated interactions with the pipeline, including login, browsing and updating the sales leads which are logged by the system; ii) the pending opportunity is prone to reach its win outcome shortly after such temporally concentrated interactions. Our model is deployed and in continual use to a large, global, B2B multinational technology enterprize (Fortune 500) with a case study. Due to the generality and flexibility of the model, it also enjoys the potential applicability to other real-world problems.

IJCAI Conference 2013 Conference Paper

Social Trust Prediction Using Rank-k Matrix Recovery

  • Jin Huang
  • Feiping Nie
  • Heng Huang
  • Yu Lei
  • Chris Ding

Trust prediction, which explores the unobserved relationships between online community users, is an emerging and important research topic in social network analysis and many web applications. Similar to other social-based recommender systems, trust relationships between users can be also modeled in the form of matrices. Recent study shows users generally establish friendship due to a few latent factors, it is therefore reasonable to assume the trust matrices are of low-rank. As a result, many recommendation system strategies can be applied here. In particular, trace norm minimization, which uses matrix’s trace norm to approximate its rank, is especially appealing. However, recent articles cast doubts on the validity of trace norm approximation. In this paper, instead of using trace norm minimization, we propose a new robust rank-k matrix completion method, which explicitly seeks a matrix with exact rank. Moreover, our method is robust to noise or corrupted observations. We optimize the new objective function in an alternative manner, based on a combination of ancillary variables and Augmented Lagrangian Multiplier (ALM) Method. We perform the experiments on three real-world data sets and all empirical results demonstrate the effectiveness of our method.

AAAI Conference 2013 Conference Paper

Spectral Rotation versus K-Means in Spectral Clustering

  • Jin Huang
  • Feiping Nie
  • Heng Huang

Spectral clustering has been a popular data clustering algorithm. This category of approaches often resort to other clustering methods, such as K-Means, to get the final cluster. The potential flaw of such common practice is that the obtained relaxed continuous spectral solution could severely deviate from the true discrete solution. In this paper, we propose to impose an additional orthonormal constraint to better approximate the optimal continuous solution to the graph cut objective functions. Such a method, called spectral rotation in literature, optimizes the spectral clustering objective functions better than K-Means, and improves the clustering accuracy. We would provide efficient algorithm to solve the new problem rigorously, which is not significantly more costly than K-Means. We also establish the connection between our method and K-Means to provide theoretical motivation of our method. Experimental results show that our algorithm consistently reaches better cut and meanwhile outperforms in clustering metrics than classic spectral clustering methods.

AAAI Conference 2013 Conference Paper

Supervised and Projected Sparse Coding for Image Classification

  • Jin Huang
  • Feiping Nie
  • Heng Huang
  • Chris Ding

Classic sparse representation for classification (SRC) method fails to incorporate the label information of training images, and meanwhile has a poor scalability due to the expensive computation for `1 norm. In this paper, we propose a novel subspace sparse coding method with utilizing label information to effectively classify the images in the subspace. Our new approach unifies the tasks of dimension reduction and supervised sparse vector learning, by simultaneously preserving the data sparse structure and meanwhile seeking the optimal projection direction in the training stage, therefore accelerates the classification process in the test stage. Our method achieves both flat and structured sparsity for the vector representations, therefore making our framework more discriminative during the subspace learning and subsequent classification. The empirical results on 4 benchmark data sets demonstrate the effectiveness of our method.

IJCAI Conference 2013 Conference Paper

Towards Effective Prioritizing Water Pipe Replacement and Rehabilitation

  • Junchi Yan
  • Yu Wang
  • Ke Zhou
  • Jin Huang
  • Chunhua Tian
  • Hongyuan Zha
  • Weishan Dong

Water pipe failures can not only have a great impact on people’s daily life but also cause significant waste of water which is an essential and precious resource to human beings. As a result, preventative maintenance for water pipes, particularly in urbanscale networks, is of great importance for a sustainable society. To achieve effective replacement and rehabilitation, failure prediction aims to proactively find those ‘most-likely-to-fail’ pipes becomes vital and has been attracting more attention from both academia and industry, especially from the civil engineering field. This paper presents an alreadydeployed industrial computational system for pipe failure prediction. As an alternative to risk matrix methods often depending on ad-hoc domain heuristics, learning based methods are adopted using the attributes with respect to physical, environmental, operational conditions and etc. Further challenge arises in practice when lacking of profile attributes. A dive into the failure records shows that the failure event sequences typically exhibit temporal clustering patterns, which motivates us to use the stochastic process to tackle the failure prediction task. Specifically, the failure sequence is formulated as a self-exciting stochastic process which is, to our best knowledge, a novel formulation for pipe failure prediction. And we show that it outperforms a baseline assuming the failure risk grows linearly with aging. Broad new problems and research points for the machine learning community are also introduced for future work.

IJCAI Conference 2007 Conference Paper

  • Jin Huang
  • Charles X. Ling

Evaluation measures play an important role in machine learning because they are used not only to compare different learning algorithms, but also often as goals to optimize in constructing learning models. Both formal and empirical work has been published in comparing evaluation measures. In this paper, we propose a general approach to construct new measures based on the existing ones, and we prove that the new measures are consistent with, and finer than, the existing ones. We also show that the new measure is more correlated to RMS (Root Mean Square error) with artificial datasets. Finally, we demonstrate experimentally that the greedy-search based algorithm (such as artificial neural networks) trained with the new and finer measure usually can achieve better prediction performance. This provides a general approach to improve the predictive performance of existing learning algorithms based on greedy search.

IJCAI Conference 2003 Conference Paper

AUC: a Statistically Consistent and more Discriminating Measure than Accuracy

  • Charles X. Ling
  • Jin Huang
  • Harry Zhang

Predictive accuracy has been used as the main and often only evaluation criterion for the predictive performance of classification learning algorithms. In recent years, the area under the ROC (Receiver Operating Characteristics) curve, or simply AUC, has been proposed as an alternative single-number measure for evaluating learning algorithms. In this paper, we prove that AUC is a better measure than accuracy. More specifically, we present rigourous definitions on consistency and discriminancy in comparing two evaluation measures for learning algorithms. We then present empirical evaluations and a formal proof to establish that AUC is indeed statistically consistent and more discriminating than accuracy. Our result is quite significant since we formally prove that, for the first time, AUC is a better measure than accuracy in the evaluation of learning algorithms.

v2026.09.13