Arrow Research search

Author name cluster

Sheng Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

58 papers
2 author rows

Possible papers

58

AAAI Conference 2026 Short Paper

Doubly Robust Causal Estimation Under Multi-View Network Interference (Student Abstract)

  • Hanzhang Yuan
  • Sheng Li

Estimating causal effects under network interference is challenging especially when edges are heterogeneous and nodes share latent dependencies. We study this realistic setting and propose MVDR, a targeted maximum likelihood (TMLE) framework that learns multi-view representations of covariates and exposure on heterogeneous networks while achieving double robustness: consistency holds if either the outcome model or the exposure density is correctly specified. MVDR supports multiple network interventions using only the observed network structure. On three semi-synthetic datasets, MVDR reduces intervention-level prediction error against baselines, and remains stable under misspecification.

EAAI Journal 2026 Journal Article

Multi-scale fusion global perception network for gastrointestinal disease classification

  • Sheng Li
  • Yulin Yu
  • Xiongxiong He

With the rising incidence of gastrointestinal diseases, improving the accuracy of automated diagnosis has become a critical research focus in medical image analysis. Efficient and precise diagnostic techniques not only advance the interdisciplinary development of computer vision and medical imaging but also play a vital role in enabling early detection and personalized treatment in clinical practice. To enhance the model’s adaptability to various gastrointestinal lesions and image quality variations while improving generalization and recognition accuracy, we propose a multi-scale fusion-based global perception network. The multi-scale cross-fusion attention module strengthens the model’s ability to detect diverse lesions and lesion areas. Meanwhile, the global perception module facilitates interactions across different regions, effectively capturing long-range dependencies to mitigate the challenges posed by high inter-class similarity and large intra-class variance. Additionally, we introduce an advanced feature fusion framework that integrates both shallow and deep features, ensuring a comprehensive utilization of image details and global context. The adaptive feature selection mechanism further enables the network to flexibly adjust to different lesion types, capturing complex gastrointestinal features and overcoming the limitations of single-scale representations. We evaluated our method on a private five-class small intestine dataset, the public five-class Kvasir-Capsule dataset, the public three-class Kvasir dataset, the public three-class Hyper-Kvasir dataset, and the public four-class Piccolo dataset. Experimental results demonstrate that our proposed method outperforms the comparative methods. The overall classification accuracies achieved were 97. 58%, 98. 33%, 97. 17%, 95. 22%, and 94. 44%, respectively. These results not only demonstrate the superiority of our method on specific datasets but also highlight its strong generalization ability across different datasets and clinical scenarios. This not only validates the practical effectiveness of the proposed model in complex clinical imaging scenarios, but also provides a solid theoretical foundation and technical support for the future design and deployment of intelligent gastrointestinal disease diagnosis systems.

EAAI Journal 2026 Journal Article

SteelFlow: A time series prediction framework for carbon content and temperature in converter steelmaking

  • DeHao Han
  • Jianzheng Zhang
  • Hongbing Wang
  • Sheng Li
  • Anjun Xu

Accurate prediction of endpoint carbon content and temperature in converter steelmaking is severely constrained by extremely sparse measurements, where only one or two offline observations are available per heat at mid-blow (TSC) and end-blow (TSO). Most existing time-series and data-driven models implicitly assume continuous target observations, an assumption violated in practical steelmaking. To address this limitation, this paper proposes SteelFlow, a mechanism-constrained time-series prediction framework that reformulates endpoint prediction as a trajectory learning problem under extreme supervision sparsity by reconstructing physically reachable carbon–temperature trajectories between TSC and TSO using metallurgical reaction constraints. By leveraging historical trend information from similar heats as guidance rather than ground truth, SteelFlow enables anticipatory endpoint prediction during the final blowing stage. To robustly model temporal dependencies under sparse supervision, a dedicated Transformer-based prediction model, termed Steelformer, is developed by jointly integrating trend-guided representations, derivative features, and a relative-time-aware attention mechanism. Experiments on a real industrial dataset consisting of 2012 heats collected from multiple top-and-bottom blown converters (70–100 t) demonstrate that the proposed approach consistently outperforms conventional endpoint prediction methods and representative deep learning baselines in terms of accuracy and stability, evaluated using both industrial tolerance criteria and standard regression metrics.

TMLR Journal 2026 Journal Article

The Clever Hans Mirage: A Comprehensive Survey on Spurious Correlations in Machine Learning

  • Wenqian Ye
  • Luyang Jiang
  • Eric Xie
  • Guangtao Zheng
  • Yunsheng Ma
  • Xu Cao
  • Dongliang Guo
  • Daiqing Qi

Back in the early 20th century, a horse named Hans appeared to perform arithmetic and other intellectual tasks during exhibitions in Germany, while it actually relied solely on involuntary cues in the body language from the human trainer. Modern machine learning models are no different. These models are known to be sensitive to spurious correlations between non-essential features of the inputs (e.g., background, texture, and secondary objects) and the corresponding labels. Such features and their correlations with the labels are known as spurious because they tend to change with shifts in real-world data distributions, which can negatively impact the model's generalization and robustness. In this paper, we provide a comprehensive survey of this emerging issue, along with a fine-grained taxonomy of existing state-of-the-art methods for addressing spurious correlations in machine learning models. Additionally, we summarize existing datasets, benchmarks, and metrics to facilitate future research. The paper concludes with a discussion of the broader impacts, the recent advancements, and future challenges in the era of generative AI, aiming to provide valuable insights for researchers in the related domains of the machine learning community.

AAAI Conference 2025 Conference Paper

Bridging Knowledge Gap Between Image Inpainting and Large-Area Visible Watermark Removal

  • Yicheng Leng
  • Chaowei Fang
  • Junye Chen
  • Yixiang Fang
  • Sheng Li
  • Guanbin Li

Visible watermark removal which involves watermark cleaning and background content restoration is pivotal to evaluate the resilience of watermarks. Existing deep neural network (DNN)-based models still struggle with large-area watermarks and are overly dependent on the quality of watermark mask prediction. To overcome these challenges, we introduce a novel feature adapting framework that leverages the representation modeling capacity of a pre-trained image inpainting model. Our approach bridges the knowledge gap between image inpainting and watermark removal by fusing information of the residual background content beneath watermarks into the inpainting backbone model. We establish a dual-branch system to capture and embed features from the residual background content, which are merged into intermediate features of the inpainting backbone model via gated feature fusion modules. Moreover, for relieving the dependence on high-quality watermark masks, we introduce a new training paradigm by utilizing coarse watermark masks to guide the inference process. This contributes to a visible image removal model which is insensitive to the quality of watermark mask during testing. Extensive experiments on both a large-scale synthesized dataset and a real-world dataset demonstrate that our approach significantly outperforms existing state-of-the-art methods. The source code is available in the supplementary materials.

AAAI Conference 2025 Conference Paper

Embedding Robust Watermarking into Pattern to Protect the Copyright of Ceramic Artifacts

  • Lei Tan
  • Yuliang Xue
  • Guobiao Li
  • Zhenxing Qian
  • Sheng Li
  • Chunlei Bao

Ceramic artworks with elegant patterns present enormous collectible value and profits. To claim the copyright, the builder usually pastes their conspicuous stamp on the bottom or side of the ceramic artworks, which inevitably affects the external image of the artwork. In addition, the stamp is weak in resisting forgery attacks due to its visible nature. To address the above issues, we propose in this paper a novel framework for embedding invisible watermarking into patterns of the ceramic artworks. In the framework, a template-based watermarking embedding scheme is designed to map the watermark to an invisible template, which is added to the ceramic pattern to create its watermarked version. A distortion layer is further proposed to model the distortion of ceramic patterns in the ceramic manufacturing process, where a color-halftoning and an adaptive brightness adjustment strategy are developed to counter the print and firing operations that introduce the most significant distortions. Finally, a deep decoder is learned to extract the watermarking from the distorted pattern. Various experiments have been conducted to demonstrate the advantage of our proposed method for protecting the copyright of the ceramic artworks, which provides reliable watermark extraction accuracy without the need for a conspicuous stamp.

NeurIPS Conference 2025 Conference Paper

Improve Temporal Reasoning in Multimodal Large Language Models via Video Contrastive Decoding

  • Daiqing Qi
  • Dongliang Guo
  • Hanzhang Yuan
  • Handong Zhao
  • Mengxuan Hu
  • Lehan Yang
  • Sheng Li

A major distinction between video and image understanding is that the former requires reasoning over time. Existing Video Large Language Models (VLLMs) demonstrate promising performance in general video understanding, such as brief captioning or object recognition within individual frames. However, they often struggle with temporal reasoning such as understanding continuous actions or tracking object transformations over time—which typically demands the integration of multiple frames in a temporally coherent manner. We first explore and explain such failures in Video LLMs from the perspective of \textit{language and ``image'' priors. } While existing research has attempted to enhance the temporal understanding of VLLMs through various training strategies, the demand for expensive computational resources and training data often presents significant barriers. To this end, we further propose a simple yet novel idea for improving temporal reasoning in videos at no additional training cost. Specifically, to better capture the temporal structure across multiple frames—the key to effective temporal reasoning—we distort the temporal consistency in key frames \textit{during the decoding phase}. Such corruption induces time-insensitive wrong responses from the model, which are then contrastively avoided when generating the final correct output. In this way, the model is encouraged to perform more temporally coherent reasoning. Our method yields consistent improvements across both temporal-specific and general video understanding benchmarks, demonstrating its effectiveness and generalizability.

IJCAI Conference 2025 Conference Paper

Large Language Models for Causal Discovery: Current Landscape and Future Directions

  • Guangya Wan
  • Yunsheng Lu
  • Yuqi Wu
  • Mengxuan Hu
  • Sheng Li

Causal discovery (CD) and Large Language Models (LLMs) have emerged as transformative fields in artificial intelligence that have evolved largely independently. While CD specializes in uncovering cause-effect relationships from data, and LLMs excel at natural language processing and generation, their integration presents unique opportunities for advancing causal understanding. This survey examines how LLMs are transforming CD across three key dimensions: direct causal extraction from text, integration of domain knowledge into statistical methods, and refinement of causal structures. We systematically analyze approaches that leverage LLMs for CD tasks, highlighting their innovative use of metadata and natural language for causal inference. Our analysis reveals both LLMs' potential to enhance traditional CD methods and their current limitations as imperfect expert systems. We identify key research gaps, outline evaluation frameworks and benchmarks for LLM-based causal discovery, and advocate future research efforts for leveraging LLMs in causality research. As the first comprehensive examination of the synergy between LLMs and CD, this work lays the groundwork for future advances in the field.

AAAI Conference 2025 Conference Paper

Physical Marker: Revealing Invisible Hyperlinks Hidden in Printed Trademarks

  • Yuliang Xue
  • Lei Tan
  • Guobiao Li
  • Zhenxing Qian
  • Sheng Li
  • Xinpeng Zhang

Embedding links in brand logos is a promising technology, which allows consumers to access the online information of products by capturing physical logo images. Previous physical data hiding methods primarily embed data within cover media in a global manner, making them ineffective for processing brand logos in vector graphics format with a transparent background. To address this issue, we propose in this paper a novel physical deep hiding scheme for invisibly embedding links in printed trademarks. Specifically, the encoder embeds links only into the area of the brand logo under the constraints of a mask, which is generated from the transparency information of the logo image. A background variation distortion is introduced into the distortion layer that approximate practical logo print-camera environments, such that the decoder could be learnt to retrieve the link from the camera-captured logo with various backgrounds. A feature prompt subspace modulator is further proposed and employed in the encoder to enhance the invisibility of the encoded logo pattern and in the decoder to boost hyperlink extraction accuracy. Various experiments have been conducted to demonstrate the advantage of our proposed method for embedding links in printed brand logos, which provides reliable extraction accuracy under both simulated and real scenarios.

AAAI Conference 2025 Conference Paper

Trustworthy AI Meets Educational Assessment: Challenges and Opportunities

  • Sheng Li

Artificial intelligence (AI) has made substantial impacts in numerous fields, including education. Within education, learning and assessment are two key areas. Although many AI techniques have been applied to improve teaching and learning, their potential in educational assessment remains underexplored. This paper explores the intersection of AI and educational assessment and presents a rich landscape of challenges and opportunities, especially in the context of trustworthy AI, including fairness, transparency, accountability, explainability, and robustness. We will begin by outlining the foundations of trustworthy AI and educational assessment. Next, we will delve into the application of trustworthy AI for various assessment tasks, such as test item generation, test design, and automated scoring. In addition, the talk will also discuss how insights from educational measurement theory, such as item response theory (IRT) and validity frameworks, can inform the development and evaluation of trustworthy AI models. These frameworks help ensure that AI systems in education are not only accurate, but also equitable and aligned with educational goals. Finally, we will highlight future research directions, focusing on the integration of ethical AI principles into educational technology and the need for interdisciplinary collaboration to tackle the emerging challenges in this field. The aim is to foster a new generation of AI-powered educational tools that are both innovative and trustworthy, ultimately contributing to a more equitable and more effective educational landscape.

AAAI Conference 2025 Conference Paper

UFID: A Unified Framework for Black-box Input-level Backdoor Detection on Diffusion Models

  • Zihan Guan
  • Mengxuan Hu
  • Sheng Li
  • Anil Kumar Vullikanti

Diffusion models are vulnerable to backdoor attacks, where malicious attackers inject backdoors by poisoning certain training samples during the training stage. This poses a significant threat to real-world applications in the Model-as-a-Service (MaaS) scenario, where users query diffusion models through APIs or directly download them from the internet. To mitigate the threat of backdoor attacks under MaaS, black-box input-level backdoor detection has drawn recent interest, where defenders aim to build a firewall that filters out backdoor samples in the inference stage, with access only to input queries and the generated results from diffusion models. Despite some preliminary explorations on the traditional classification tasks, these methods cannot be directly applied to the generative tasks due to two major challenges: (1) more diverse failures and (2) a multi-modality attack surface. In this paper, we propose a black-box input-level backdoor detection framework on diffusion models, called UFID. Our defense is motivated by an insightful causal analysis: Backdoor attacks serve as the confounder, introducing a spurious path from input to target images, which remains consistent even when we perturb the input samples with Gaussian noise. We further validate the intuition with theoretical analysis. Extensive experiments across different datasets on both conditional and unconditional diffusion models show that our method achieves superb performance on detection effectiveness and run-time efficiency.

AAAI Conference 2024 Short Paper

BadSAM: Exploring Security Vulnerabilities of SAM via Backdoor Attacks (Student Abstract)

  • Zihan Guan
  • Mengxuan Hu
  • Zhongliang Zhou
  • Jielu Zhang
  • Sheng Li
  • Ninghao Liu

Image segmentation is foundational to computer vision applications, and the Segment Anything Model (SAM) has become a leading base model for these tasks. However, SAM falters in specialized downstream challenges, leading to various customized SAM models. We introduce BadSAM, a backdoor attack tailored for SAM, revealing that customized models can harbor malicious behaviors. Using the CAMO dataset, we confirm BadSAM's efficacy and identify SAM vulnerabilities. This study paves the way for the development of more secure and customizable vision foundation models.

TMLR Journal 2024 Journal Article

BBCaL: Black-box Backdoor Detection under the Causality Lens

  • Mengxuan Hu
  • Zihan Guan
  • Junfeng Guo
  • Zhongliang Zhou
  • Jielu Zhang
  • Sheng Li

Deep Neural Networks (DNNs) are known to be vulnerable to backdoor attacks, where attackers can inject hidden backdoors during the training stage. This poses a serious threat to the Model-as-a-Service setting, where downstream users directly utilize third-party models (e.g., HuggingFace Hub, ChatGPT). To this end, we study the inference-stage black-box backdoor detection problem in the paper, where defenders aim to build a firewall to filter out the backdoor inputs in the inference stage, with only input samples and prediction labels available. Existing investigations on this problem either rely on strong assumptions on types of triggers and attacks or suffer from poor efficiency. To build a more generalized and efficient method, we first provide a novel causality-based lens to analyze heterogeneous prediction behaviors for clean and backdoored samples in the inference stage, considering both sample-specific and sample-agnostic backdoor attacks. Motivated by the causal analysis and do-calculus in causal inference, we introduce Black-box Backdoor detection under the Causality Lens (BBCaL) which distinguishes backdoor and clean samples by analyzing prediction consistency after progressively constructing counterfactual samples. Theoretical analysis also sheds light on the effectiveness of the BBCaL. Extensive experiments on three benchmark datasets validate the effectiveness and efficiency of our method.

NeurIPS Conference 2024 Conference Paper

Disentangled Style Domain for Implicit $z$-Watermark Towards Copyright Protection

  • Junqiang Huang
  • Zhaojun Guo
  • Ge Luo
  • Zhenxing Qian
  • Sheng Li
  • Xinpeng Zhang

Text-to-image models have shown surprising performance in high-quality image generation, while also raising intensified concerns about the unauthorized usage of personal dataset in training and personalized fine-tuning. Recent approaches, embedding watermarks, introducing perturbations, and inserting backdoors into datasets, rely on adding minor information vulnerable to adversarial training, limiting their ability to detect unauthorized data usage. In this paper, we introduce a novel implicit Zero-Watermarking scheme that first utilizes the disentangled style domain to detect unauthorized dataset usage in text-to-image models. Specifically, our approach generates the watermark from the disentangled style domain, enabling self-generalization and mutual exclusivity within the style domain anchored by protected units. The domain achieves the maximum concealed offset of probability distribution through both the injection of identifier $z$ and dynamic contrastive learning, facilitating the structured delineation of dataset copyright boundaries for multiple sources of styles and contents. Additionally, we introduce the concept of watermark distribution to establish a verification mechanism for copyright ownership of hybrid or partial infringements, addressing deficiencies in the traditional mechanism of dataset copyright ownership for AI mimicry. Notably, our method achieves one-sample verification for copyright ownership in AI mimic generations. The code is available at: [https: //github. com/Hlufies/ZWatermarking](https: //github. com/Hlufies/ZWatermarking)

TMLR Journal 2024 Journal Article

DTRNet: Precisely Correcting Selection Bias in Individual-Level Continuous Treatment Effect Estimation by Reweighted Disentangled Representation

  • Mengxuan Hu
  • Zhixuan Chu
  • Sheng Li

Estimating the individual-level continuous treatment effect holds significant practical importance in various decision-making domains, such as personalized healthcare and customized marketing. However, most current methods for individual treatment effect estimation are limited to discrete treatments and struggle to precisely adjust for selection bias under continuous settings, leading to inaccurate estimation. To address these challenges, we propose a novel Disentangled Representation Network (DTRNet) to estimate the individualized dose-response function (IDRF), which learns disentangled representations and precisely adjusts for selection bias. To the best of our knowledge, our work is the first attempt to precisely adjust for selection bias in continuous settings. Extensive results on synthetic and semi-synthetic datasets demonstrate that our DTRNet outperforms most state-of-the-art methods. Our code is available at \href{https://github.com/xuanxuan03021/DTRNet_final_2}{DTRNet}.

TMLR Journal 2024 Journal Article

Dual-windowed Vision Transformer with Angular Self- Attention

  • Weili Shi
  • Sheng Li

Following the great success in natural language processing, transformer-based models have emerged as the competitive model against the convolutional neural networks in computer vision. Vision transformer (ViT) and its subsequent variants have exhibited promising performance in tasks such as image classification, object detection and semantic segmentation. The core of vision transformers is the self-attention mechanism, which models the long-range dependency of different tokens. Conventionally, the attention matrix in self-attention is calculated by the scaled dot-product of \textit{query} (Q) and \textit{key} (K). In this case, the attention weight would depend on norm of Q and K as well as the angle between them. In this paper, we propose a new attention mechanism named angular self-attention, which replaces the scaled dot-product operation with the angular function in order to effectively model the relationship between tokens. In particular, we propose two forms of functions: quadratic and cosine functions, for our angular self-attention. Based on angular self-attention, we design a new vision transformer architecture called dual-windowed angular vision transformer (\textbf{DWAViT}). DWAViT is a hierarchical-structured model characterized by the angular self-attention and a new local window mechanism. We evaluate DWAViT on multiple computer vision benchmarks, including image classification on ImageNet-1K, object detection on COCO, and semantic segmentation on ADE20K. Our experimental results also suggest that our model can achieve promising performance on the tasks while maintaining comparable computational cost with that of the baseline models (e.g., Swin Transformer).

NeurIPS Conference 2024 Conference Paper

Easy Regional Contrastive Learning of Expressive Fashion Representations

  • Daiqing Qi
  • Handong Zhao
  • Sheng Li

When learning vision-language models (VLM) for the fashion domain, most existing works design new architectures from vanilla BERT with additional objectives, or perform dense multi-task learning with fashion-specific tasks. Though progress has been made, their architecture or objectives are often intricate and the extendibility is limited. By contrast, with simple architecture (comprising only two unimodal encoders) and just the contrastive objective, popular pre-trained VL models (e. g. , CLIP) achieve superior performance in general domains, which are further easily extended to downstream tasks. However, inheriting such benefits of CLIP in the fashion domain is non-trivial in the presence of the notable domain gap. Empirically, we find that directly finetuning on fashion data leads CLIP to frequently ignore minor yet important details such as logos and composition, which are critical in fashion tasks such as retrieval and captioning. In this work, to maintain CLIP's simple architecture and objective while explicitly attending to fashion details, we propose $E^2$: Easy Regional Contrastive Learning of Expressive Fashion Representations. $E^2$ introduces only a few selection tokens and fusion blocks (just 1. 9\% additional parameters in total) with only contrastive losses. Despite lightweight, in our primary focus, cross-modal retrieval, $E^2$ notably outperforms existing fashion VLMs with various fashion-specific objectives. Moreover, thanks to CLIP's widespread use in downstream tasks in general domains (e. g. , zero-shot composed image retrieval and image captioning), our model can easily extend these models from general domain to the fashion domain with notable improvement. To conduct a comprehensive evaluation, we further collect data from Amazon Reviews to build a new dataset for cross-modal retrieval in the fashion domain.

AAAI Conference 2024 Conference Paper

Frozen CLIP Transformer Is an Efficient Point Cloud Encoder

  • Xiaoshui Huang
  • Zhou Huang
  • Sheng Li
  • Wentao Qu
  • Tong He
  • Yuenan Hou
  • Yifan Zuo
  • Wanli Ouyang

The pretrain-finetune paradigm has achieved great success in NLP and 2D image fields because of the high-quality representation ability and transferability of their pretrained models. However, pretraining such a strong model is difficult in the 3D point cloud field due to the limited amount of point cloud sequences. This paper introduces Efficient Point Cloud Learning (EPCL), an effective and efficient point cloud learner for directly training high-quality point cloud models with a frozen CLIP transformer. Our EPCL connects the 2D and 3D modalities by semantically aligning the image features and point cloud features without paired 2D-3D data. Specifically, the input point cloud is divided into a series of local patches, which are converted to token embeddings by the designed point cloud tokenizer. These token embeddings are concatenated with a task token and fed into the frozen CLIP transformer to learn point cloud representation. The intuition is that the proposed point cloud tokenizer projects the input point cloud into a unified token space that is similar to the 2D images. Comprehensive experiments on 3D detection, semantic segmentation, classification and few-shot learning demonstrate that the CLIP transformer can serve as an efficient point cloud encoder and our method achieves promising performance on both indoor and outdoor benchmarks. In particular, performance gains brought by our EPCL are 19.7 AP50 on ScanNet V2 detection, 4.4 mIoU on S3DIS segmentation and 1.2 mIoU on SemanticKITTI segmentation compared to contemporary pretrained models. Code is available at \url{https://github.com/XiaoshuiHuang/EPCL}.

AAAI Conference 2024 Conference Paper

LLMRG: Improving Recommendations through Large Language Model Reasoning Graphs

  • Yan Wang
  • Zhixuan Chu
  • Xin Ouyang
  • Simeng Wang
  • Hongyan Hao
  • Yue Shen
  • Jinjie Gu
  • Siqiao Xue

Recommendation systems aim to provide users with relevant suggestions, but often lack interpretability and fail to capture higher-level semantic relationships between user behaviors and profiles. In this paper, we propose a novel approach that leverages large language models (LLMs) to construct personalized reasoning graphs. These graphs link a user's profile and behavioral sequences through causal and logical inferences, representing the user's interests in an interpretable way. Our approach, LLM reasoning graphs (LLMRG), has four components: chained graph reasoning, divergent extension, self-verification and scoring, and knowledge base self-improvement. The resulting reasoning graph is encoded using graph neural networks, which serves as additional input to improve conventional recommender systems, without requiring extra user or item information. Our approach demonstrates how LLMs can enable more logical and interpretable recommender systems through personalized reasoning graphs. LLMRG allows recommendations to benefit from both engineered recommendation systems and LLM-derived reasoning graphs. We demonstrate the effectiveness of LLMRG on benchmarks and real-world scenarios in enhancing base recommendation models.

AAAI Conference 2024 Conference Paper

Open-Set Graph Domain Adaptation via Separate Domain Alignment

  • Yu Wang
  • Ronghang Zhu
  • Pengsheng Ji
  • Sheng Li

Domain adaptation has become an attractive learning paradigm, as it can leverage source domains with rich labels to deal with classification tasks in an unlabeled target domain. A few recent studies develop domain adaptation approaches for graph-structured data. In the case of node classification task, current domain adaptation methods only focus on the closed-set setting, where source and target domains share the same label space. A more practical assumption is that the target domain may contain new classes that are not included in the source domain. Therefore, in this paper, we introduce a novel and challenging problem for graphs, i.e., open-set domain adaptive node classification, and propose a new approach to solve it. Specifically, we develop an algorithm for efficient knowledge transfer from a labeled source graph to an unlabeled target graph under a separate domain alignment (SDA) strategy, in order to learn discriminative feature representations for the target graph. Our goal is to not only correctly classify target nodes into the known classes, but also classify unseen types of nodes into an unknown class. Experimental results on real-world datasets show that our method outperforms existing methods on graph domain adaptation.

TMLR Journal 2024 Journal Article

Revealing an Overlooked Challenge in Class-Incremental Graph Learning

  • Daiqing Qi
  • Handong Zhao
  • Xiaowei Jia
  • Sheng Li

Graph Neural Networks (GNNs), which effectively learn from static graph-structured data, become ineffective when directly applied to streaming data in a continual learning (CL) scenario. In CL, historical data are not available during the current stage due to a number of reasons, such as limited storage, GDPR1 data retention policy, to name a few. A few recent works study this problem, however, they overlook the uniqueness of continual graph learning (CGL), compared to well-studied continual image classification: the unavailability of previous training data further poses challenges to inference in CGL, in additional to the well-known catastrophic forgetting problem. While existing works make a strong assumption that full access of historical data is unavailable during training but provided during inference, which potentially contradicts the continual learning paradigm Van de Ven & Tolias (2019), we study continual graph learning without this strong and contradictory assumption. In this case, without being re-inserted into previous training graphs for inference, streaming test nodes are often more sparsely connected, which makes the inference more difficult due to insufficient neighborhood information. In this work, we propose ReplayGNN (ReGNN) to jointly solve the above two challenges without memory buffers: catastrophic forgetting and poor neighbor information during inference. Extensive experiments demonstrate the effectiveness of our model over baseline models and its effectiveness in different cases with different levels of neighbor information available.

AAAI Conference 2024 Conference Paper

Task-Driven Causal Feature Distillation: Towards Trustworthy Risk Prediction

  • Zhixuan Chu
  • Mengxuan Hu
  • Qing Cui
  • Longfei Li
  • Sheng Li

Since artificial intelligence has seen tremendous recent successes in many areas, it has sparked great interest in its potential for trustworthy and interpretable risk prediction. However, most models lack causal reasoning and struggle with class imbalance, leading to poor precision and recall. To address this, we propose a Task-Driven Causal Feature Distillation model (TDCFD) to transform original feature values into causal feature attributions for the specific risk prediction task. The causal feature attribution helps describe how much contribution the value of this feature can make to the risk prediction result. After the causal feature distillation, a deep neural network is applied to produce trustworthy prediction results with causal interpretability and high precision/recall. We evaluate the performance of our TDCFD method on several synthetic and real datasets, and the results demonstrate its superiority over the state-of-the-art methods regarding precision, recall, interpretability, and causality.

IJCAI Conference 2023 Conference Paper

Graph-based Semi-supervised Local Clustering with Few Labeled Nodes

  • Zhaiming Shen
  • Ming-Jun Lai
  • Sheng Li

Local clustering aims at extracting a local structure inside a graph without the necessity of knowing the entire graph structure. As the local structure is usually small in size compared to the entire graph, one can think of it as a compressive sensing problem where the indices of target cluster can be thought as a sparse solution to a linear system. In this paper, we apply this idea based on two pioneering works under the same framework and propose a new semi-supervised local clustering approach using only few labeled nodes. Our approach improves the existing works by making the initial cut to be the entire graph and hence overcomes a major limitation of the existing works, which is the low quality of initial cut. Extensive experimental results on various datasets demonstrate the effectiveness of our approach.

IJCAI Conference 2023 Conference Paper

pTSE: A Multi-model Ensemble Method for Probabilistic Time Series Forecasting

  • Yunyi Zhou
  • Zhixuan Chu
  • Yijia Ruan
  • Ge Jin
  • Yuchen Huang
  • Sheng Li

Various probabilistic time series forecasting models have sprung up and shown remarkably good performance. However, the choice of model highly relies on the characteristics of the input time series and the fixed distribution that model is based on. Due to the fact that the probability distributions cannot be averaged over different models straightforwardly, the current time series model ensemble methods cannot be directly applied to improve the robustness and accuracy of forecasting. To address this issue, we propose pTSE, a multi-model distribution ensemble method for probabilistic forecasting based on Hidden Markov Model (HMM). pTSE only takes off-the-shelf outputs from member models without requiring further information about each model. Besides, we provide a complete theoretical analysis of pTSE to prove that the empirical distribution of time series subject to an HMM will converge to the stationary distribution almost surely. Experiments on benchmarks show the superiority of pTSE over all member models and competitive ensemble methods.

TMLR Journal 2023 Journal Article

Semi-Supervised Single Domain Generalization with Label-Free Adversarial Data Augmentation

  • Ronghang Zhu
  • Xiang Yu
  • Sheng Li

Domain generalization (DG) has attracted increasing attention recently, as it seeks to improve the generalization ability of visual recognition models to unseen target domains. DG leverages multiple source domains for model training, while single domain generalization (SDG) further restricts such setting by exploiting only a single source domain. Nevertheless, both DG and SDG assume that the source domains are fully labeled, which might not be practical in many real world scenarios. In this paper, we present a new problem, i.e., semi-supervised single domain generalization (SS-SDG), which aims to train a model with a partially labeled single source domain to generalize to multiple unseen testing domains. We propose an effective framework to address this problem. In particular, we design a label-free adversarial data augmentation strategy to diversify the source domain, and propose a novel multi-pair FixMatch loss to generalize classifiers to unseen testing domains. Extensive experiments on OfficeHome, PACS and DomainNet20 datasets show that our method surpasses the latest SDG and semi-supervised methods. Moreover, on PACS and DomainNet20, our method approaches the fully supervised ERM upper bound within $5\%$ gap, but only uses less than $8\%$ of the labels.

AAAI Conference 2023 Conference Paper

Steganography of Steganographic Networks

  • Guobiao Li
  • Sheng Li
  • Meiling Li
  • Xinpeng Zhang
  • Zhenxing Qian

Steganography is a technique for covert communication between two parties. With the rapid development of deep neural networks (DNN), more and more steganographic networks are proposed recently, which are shown to be promising to achieve good performance. Unlike the traditional handcrafted steganographic tools, a steganographic network is relatively large in size. It raises concerns on how to covertly transmit the steganographic network in public channels, which is a crucial stage in the pipeline of steganography in real world applications. To address such an issue, we propose a novel scheme for steganography of steganographic networks in this paper. Unlike the existing steganographic schemes which focus on the subtle modification of the cover data to accommodate the secrets. We propose to disguise a steganographic network (termed as the secret DNN model) into a stego DNN model which performs an ordinary machine learning task (termed as the stego task). During the model disguising, we select and tune a subset of filters in the secret DNN model to preserve its function on the secret task, where the remaining filters are reactivated according to a partial optimization strategy to disguise the whole secret DNN model into a stego DNN model. The secret DNN model can be recovered from the stego DNN model when needed. Various experiments have been conducted to demonstrate the advantage of our proposed method for covert communication of steganographic networks as well as general DNN models.

IJCAI Conference 2022 Conference Paper

AgriBERT: Knowledge-Infused Agricultural Language Models for Matching Food and Nutrition

  • Saed Rezayi
  • Zhengliang Liu
  • Zihao Wu
  • Chandra Dhakal
  • Bao Ge
  • Chen Zhen
  • Tianming Liu
  • Sheng Li

Pretraining domain-specific language models remains an important challenge which limits their applicability in various areas such as agriculture. This paper investigates the effectiveness of leveraging food related text corpora (e. g. , food and agricultural literature) in pretraining transformer-based language models. We evaluate our trained language model, called AgriBERT, on the task of semantic matching, i. e. , establishing mapping between food descriptions and nutrition data, which is a long-standing challenge in the agricultural domain. In particular, we formulate the task as an answer selection problem, fine-tune the trained language model with the help of an external source of knowledge (e. g. , FoodOn ontology), and establish a baseline for this task. The experimental results reveal that our language model substantially outperforms other language models and baselines in the task of matching food description and nutrition.

NeurIPS Conference 2022 Conference Paper

Layer Freezing & Data Sieving: Missing Pieces of a Generic Framework for Sparse Training

  • Geng Yuan
  • Yanyu Li
  • Sheng Li
  • Zhenglun Kong
  • Sergey Tulyakov
  • Xulong Tang
  • Yanzhi Wang
  • Jian Ren

Recently, sparse training has emerged as a promising paradigm for efficient deep learning on edge devices. The current research mainly devotes the efforts to reducing training costs by further increasing model sparsity. However, increasing sparsity is not always ideal since it will inevitably introduce severe accuracy degradation at an extremely high sparsity level. This paper intends to explore other possible directions to effectively and efficiently reduce sparse training costs while preserving accuracy. To this end, we investigate two techniques, namely, layer freezing and data sieving. First, the layer freezing approach has shown its success in dense model training and fine-tuning, yet it has never been adopted in the sparse training domain. Nevertheless, the unique characteristics of sparse training may hinder the incorporation of layer freezing techniques. Therefore, we analyze the feasibility and potentiality of using the layer freezing technique in sparse training and find it has the potential to save considerable training costs. Second, we propose a data sieving method for dataset-efficient training, which further reduces training costs by ensuring only a partial dataset is used throughout the entire training process. We show that both techniques can be well incorporated into the sparse training algorithm to form a generic framework, which we dub SpFDE. Our extensive experiments demonstrate that SpFDE can significantly reduce training costs while preserving accuracy from three dimensions: weight sparsity, layer freezing, and dataset sieving. Our code and models will be released.

ICRA Conference 2022 Conference Paper

Learning Emergent Discrete Message Communication for Cooperative Reinforcement Learning

  • Sheng Li
  • Yutai Zhou
  • Ross E. Allen
  • Mykel J. Kochenderfer

Communication is an important factor that en-ables agents to work cooperatively in multi-agent reinforcement learning (MARL) contexts. Prior work used continuous message communication whose high representational capacity comes at the expense of interpretability. Allowing agents to learn their own discrete emergent message communication protocols can increase the interpretability for human designers and other agents. This paper proposes a method to generate discrete messages analogous to human languages. Discrete message communication is achieved by a broadcast-and-listen mecha-nism based on self-attention. We show that discrete message communication has performance comparable to continuous message communication but with a much smaller vocabulary size. Discrete message communication protocols can potentially be used for human-agent interaction.

AAAI Conference 2022 Conference Paper

Patch Diffusion: A General Module for Face Manipulation Detection

  • Baogen Zhang
  • Sheng Li
  • Guorui Feng
  • Zhenxing Qian
  • Xinpeng Zhang

Detection of manipulated face images has attracted a lot of interest recently. Various schemes have been proposed to tackle this challenging problem, where the patch-based approaches are shown to be promising. However, the existing patch-based approaches tend to treat different patches equally, which do not fully exploit the patch discrepancy for effective feature learning. In this paper, we propose a Patch Diffusion (PD) module which can be integrated into the existing face manipulation detection networks to boost the performance. The PD consists of Discrepancy Patch Feature Learning (DPFL) and Attention-Aware Message Passing (AMP). The DPFL effectively learns the patch features by a newly designed Pairwise Patch Loss (PPLoss), which takes both the patch importance and correlations into consideration. The AMP diffuses the patches through attention-aware message passing in a graph network, where the attentions are explicitly computed based on the patch features learnt in DPFL. We integrate our PD module into four recent face manipulation detection networks, and carry out the experiments on four popular datasets. The results demonstrate that our PD module is able to boost the performance of the existing networks for face manipulation detection.

AAAI Conference 2022 Short Paper

XDC: Adversarial Adaptive Cross Domain Face Clustering (Student Abstract)

  • Saed Rezayi
  • Handong Zhao
  • Sheng Li

In this work we propose a scheme, called XDC, that uses adversarial learning to train an adaptive cross domain clustering model. XDC trains a classifier on a labeled dataset and assigns labels to an unlabeled dataset. We benefit from adversarial learning such that the target dataset takes part in the training. We also use an existing image classifiers in a plugand-play fashion (i. e. , it can be replaced with any other image classifier). Unlike existing works we update the parameters of the encoder and expose the target dataset to the model during training. We apply our model on two face dataset and one non-face dataset and obtain comparable results with state-ofthe-art face clustering models.

AAAI Conference 2021 Conference Paper

Correlative Channel-Aware Fusion for Multi-View Time Series Classification

  • Yue Bai
  • Lichen Wang
  • Zhiqiang Tao
  • Sheng Li
  • Yun Fu

Multi-view time series classification (MVTSC) aims to improve the performance by fusing the distinctive temporal information from multiple views. Existing methods for MVTSC mainly aim to fuse multi-view information at an early stage, e. g. , by extracting a common feature subspace among multiple views. However, these approaches may not fully explore the unique temporal patterns of each view in complicated time series. Additionally, the label correlations of multiple views, which are critical to boosting, are usually under-explored for the MVTSC problem. To address the aforementioned issues, we propose a Correlative Channel- Aware Fusion (C2 AF) network. First, C2 AF extracts comprehensive and robust temporal patterns by a two-stream structured encoder for each view, and derives the intra-view/interview label correlations with a concise correlation matrix. Second, a channel-aware learnable fusion mechanism is implemented through CNN to further explore the global correlative patterns. Our C2 AF is an end-to-end framework for MVTSC. Extensive experimental results on three real-world datasets demonstrate the superiority of our C2 AF over the state-ofthe-art methods. A detailed ablation study is also provided to illustrate the indispensability of each model component.

AAMAS Conference 2021 Conference Paper

Deep Implicit Coordination Graphs for Multi-agent Reinforcement Learning

  • Sheng Li
  • Jayesh K. Gupta
  • Peter Morales
  • Ross Allen
  • Mykel J. Kochenderfer

Multi-agent reinforcement learning (MARL) requires coordination to efficiently solve certain tasks. Fully centralized control is often infeasible in such domains due to the size of joint action spaces. Coordination graph formalizations allow reasoning about the joint action based on the structure of interactions. However, they often require domain expertise in their design and can be difficult for dynamic environments with changing coordination requirements. This paper introduces the deep implicit coordination graph (DICG) architecture for such scenarios. DICG consists of a module for inferring the dynamic coordination graph structure which is then used by a graph neural network module to learn to implicitly reason about the joint actions or values. DICG allows learning the tradeoff between full centralization and decentralization via standard actor-critic methods to significantly improve coordination for domains with large number of agents. We apply DICG to both centralized-training-centralized-execution and centralized-trainingdecentralized-execution regimes. We demonstrate that DICG solves the relative overgeneralization pathology in predatory-prey tasks as well as outperforms various MARL baselines on the challenging StarCraft II Multi-agent Challenge (SMAC) and traffic junction environments.

IJCAI Conference 2020 Conference Paper

A Survey on Representation Learning for User Modeling

  • Sheng Li
  • Handong Zhao

Artificial intelligent systems are changing every aspect of our daily life. In the past decades, numerous approaches have been developed to characterize user behavior, in order to deliver personalized experience to users in scenarios like online shopping or movie recommendation. This paper presents a comprehensive survey of recent advances in user modeling from the perspective of representation learning. In particular, we formulate user modeling as a process of learning latent representations for users. We discuss both the static and sequential representation learning methods for the purpose of user modeling, and review representative approaches in each category, such as matrix factorization, deep collaborative filtering, and recurrent neural networks. Both shallow and deep learning methods are reviewed and discussed. Finally, we conclude this survey and discuss a number of open research problems that would inspire further research in this field.

YNICL Journal 2020 Journal Article

Theta oscillations in prolactinomas: Neurocognitive deficits in executive controls

  • Chenglong Cao
  • Wen Wen
  • Binbin Liu
  • Pan Ma
  • Sheng Li
  • Guozheng Xu
  • Jian Song

Impairment of cognitive functions has been reported in prolactinomas. However, the electrophysiological mechanisms of response activation and response inhibition in prolactinomas remain unclear. We recorded participants' scalp electroencephalography (EEG) in a visual Go/Nogo task. Compared to the healthy controls (HCs), the patients demonstrated worse performance and their prolactin (PRL) levels negatively correlated with behavioral results. Meanwhile, patients' P300 amplitudes in the Go and Nogo conditions were smaller than the HCs. The amplitudes of N200nogo in patients were smaller than the HCs as well. Lower frontal theta power was found in the patients than the HCs in both Go and Nogo conditions, which indicated a deficit in response activation and inhibition. Moreover, the PRL levels mediated the relationship between frontal theta power and behavior performance, implying that lower frontal theta power caused the dysfunction of response control by abnormally high PRL levels. Patients also showed lower occipital alpha power than the HCs, which suggested that the impaired response inhibition may arise from deficient attention control. Taken together, the present study revealed the neurocognitive discrepancies between prolactinomas and the HCs. The frontal theta oscillation was highlighted as the electrophysiological markers of the impaired response control in prolactinomas.

IJCAI Conference 2019 Conference Paper

CensNet: Convolution with Edge-Node Switching in Graph Neural Networks

  • Xiaodong Jiang
  • Pengsheng Ji
  • Sheng Li

In this paper, we present CensNet, Convolution with Edge-Node Switching graph neural network, for semi-supervised classification and regression in graph-structured data with both node and edge features. CensNet is a general graph embedding framework, which embeds both nodes and edges to a latent feature space. By using line graph of the original undirected graph, the role of nodes and edges are switched, and two novel graph convolution operations are proposed for feature propagation. Experimental results on real-world academic citation networks and quantum chemistry graphs show that our approach has achieved or matched the state-of-the-art performance.

IJCAI Conference 2019 Conference Paper

On the Estimation of Treatment Effect with Text Covariates

  • Liuyi Yao
  • Sheng Li
  • Yaliang Li
  • Hongfei Xue
  • Jing Gao
  • Aidong Zhang

Estimating the treatment effect benefits decision making in various domains as it can provide the potential outcomes of different choices. Existing work mainly focuses on covariates with numerical values, while how to handle covariates with textual information for treatment effect estimation is still an open question. One major challenge is how to filter out the nearly instrumental variables which are the variables more predictive to the treatment than the outcome. Conditioning on those variables to estimate the treatment effect would amplify the estimation bias. To address this challenge, we propose a conditional treatment-adversarial learning based matching method (CTAM). CTAM incorporates the treatment-adversarial learning to filter out the information related to nearly instrumental variables when learning the representations, and then it performs matching among the learned representations to estimate the treatment effects. The conditional treatment-adversarial learning helps reduce the bias of treatment effect estimation, which is demonstrated by our experimental results on both semi-synthetic and real-world datasets.

AAAI Conference 2019 Conference Paper

SADIH: Semantic-Aware DIscrete Hashing

  • Zheng Zhang
  • Guo-Sen Xie
  • Yang Li
  • Sheng Li
  • Zi Huang

Due to its low storage cost and fast query speed, hashing has been recognized to accomplish similarity search in largescale multimedia retrieval applications. Particularly, supervised hashing has recently received considerable research attention by leveraging the label information to preserve the pairwise similarities of data points in the Hamming space. However, there still remain two crucial bottlenecks: 1) the learning process of the full pairwise similarity preservation is computationally unaffordable and unscalable to deal with big data; 2) the available category information of data are not well-explored to learn discriminative hash functions. To overcome these challenges, we propose a unified Semantic- Aware DIscrete Hashing (SADIH) framework, which aims to directly embed the transformed semantic information into the asymmetric similarity approximation and discriminative hashing function learning. Specifically, a semantic-aware latent embedding is introduced to asymmetrically preserve the full pairwise similarities while skillfully handle the cumbersome n × n pairwise similarity matrix. Meanwhile, a semantic-aware autoencoder is developed to jointly preserve the data structures in the discriminative latent semantic space and perform data reconstruction. Moreover, an efficient alternating optimization algorithm is proposed to solve the resulting discrete optimization problem. Extensive experimental results on multiple large-scale datasets demonstrate that our SADIH can clearly outperform the state-of-the-art baselines with the additional benefit of lower computational costs.

IJCAI Conference 2019 Conference Paper

Scalable Block-Diagonal Locality-Constrained Projective Dictionary Learning

  • Zhao Zhang
  • Weiming Jiang
  • Zheng Zhang
  • Sheng Li
  • Guangcan Liu
  • Jie Qin

We propose a novel structured discriminative block- diagonal dictionary learning method, referred to as scalable Locality-Constrained Projective Dictionary Learning (LC-PDL), for efficient representation and classification. To improve the scalability by saving both training and testing time, our LC-PDL aims at learning a structured discriminative dictionary and a block-diagonal representation without using costly l0/l1-norm. Besides, it avoids extra time-consuming sparse reconstruction process with the well-trained dictionary for new sample as many existing models. More importantly, LC-PDL avoids using the com- plementary data matrix to learn the sub-dictionary over each class. To enhance the performance, we incorporate a locality constraint of atoms into the DL procedures to keep local information and obtain the codes of samples over each class separately. A block-diagonal discriminative approximation term is also derived to learn a discriminative projection to bridge data with their codes by extracting the special block-diagonal features from data, which can ensure the approximate coefficients to associate with its label information clearly. Then, a robust multiclass classifier is trained over extracted block-diagonal codes for accurate label predictions. Experimental results verify the effectiveness of our algorithm.

IJCAI Conference 2019 Conference Paper

Self-attentive Biaffine Dependency Parsing

  • Ying Li
  • Zhenghua Li
  • Min Zhang
  • Rui Wang
  • Sheng Li
  • Luo Si

The current state-of-the-art dependency parsing approaches employ BiLSTMs to encode input sentences. Motivated by the success of the transformer-based machine translation, this work for the first time applies the self-attention mechanism to dependency parsing as the replacement of the BiLSTM-based encoders, leading to competitive performance on both English and Chinese benchmark data. Based on the detailed error analysis, we then combine the power of both BiLSTM and self-attention via model ensembles, demonstrating their complementary capability of capturing contextual information. Finally, we explore the recently proposed contextualized word representations as extra input features, and further improve the parsing performance.

AAAI Conference 2018 Conference Paper

A Multi-Task Learning Approach for Improving Product Title Compression with User Search Log Data

  • Jingang Wang
  • Junfeng Tian
  • Long Qiu
  • Sheng Li
  • Jun Lang
  • Luo Si
  • Man Lan

It is a challenging and practical research problem to obtain effective compression of lengthy product titles for Ecommerce. This is particularly important as more and more users browse mobile E-commerce apps and more merchants make the original product titles redundant and lengthy for Search Engine Optimization. Traditional text summarization approaches often require a large amount of preprocessing costs and do not capture the important issue of conversion rate in E-commerce. This paper proposes a novel multi-task learning approach for improving product title compression with user search log data. In particular, a pointer network-based sequence-to-sequence approach is utilized for title compression with an attentive mechanism as an extractive method and an attentive encoder-decoder approach is utilized for generating user search queries. The encoding parameters (i. e. , semantic embedding of original titles) are shared among the two tasks and the attention distributions are jointly optimized. An extensive set of experiments with both human annotated data and online deployment demonstrate the advantage of the proposed research for both compression qualities and online business values.

AAAI Conference 2018 Short Paper

Contextual Collaborative Filtering for Student Response Prediction in Mixed-Format Tests

  • Shumin Jing
  • Sheng Li

The purpose of this study is to design a machine learning approach to predict the student response in mixed-format tests. Particularly, a novel contextual collaborative filtering model is proposed to extract latent factors for students and test items, by exploiting the item information. Empirical results from a simulation study validate the effectiveness of the proposed method.

AAAI Conference 2018 Conference Paper

Discriminative Semi-Coupled Projective Dictionary Learning for Low-Resolution Person Re-Identification

  • Kai Li
  • Zhengming Ding
  • Sheng Li
  • Yun Fu

Person re-identification (re-ID) is a fundamental task in automated video surveillance. In real-world visual surveillance systems, a person is often captured in quite low resolutions. So we often need to perform low-resolution person re-ID, where images captured by different cameras have great resolution divergences. Existing methods cope problem via some complicated and time-consuming strategies, making them less favorable in practice, and their performances are far from satisfactory. In this paper, we design a novel Discriminative Semi-coupled Projective Dictionary Learning (DSPDL) model to effectively and efficiently solve this problem. Specifically, we propose to jointly learn a pair of dictionaries and a mapping to bridge the gap across low(er) and high(er) resolution person images. Besides, we develop a novel graph regularizer to incorporate positive and negative image pair information in a parameterless fashion. Meanwhile, we adopt the efficient and powerful projective dictionary learning technique to boost the our efficiency. Experiments on three public datasets show the superiority of the proposed method to the state-of-the-art ones.

AAAI Conference 2018 Conference Paper

Latent Discriminant Subspace Representations for Multi-View Outlier Detection

  • Kai Li
  • Sheng Li
  • Zhengming Ding
  • Weidong Zhang
  • Yun Fu

Identifying multi-view outliers is challenging because of the complex data distributions across different views. Existing methods cope this problem by exploiting pairwise constraints across different views to obtain new feature representations, based on which certain outlier score measurements are de- fined. Due to the use of pairwise constraint, it is complicated and time-consuming for existing methods to detect outliers from three or more views. In this paper, we propose a novel method capable of detecting outliers from any number of data views. Our method first learns latent discriminant representations for all view data and defines a novel outlier score function based on the latent discriminant representations. Specifically, we represent multi-view data by a global low-rank representation shared by all views and residual representations specific to each view. Through analyzing the view-specific residual representations of all views, we can get the outlier score for every sample. Moreover, we raise the problem of detecting a third type of multi-view outliers which are neglected by existing methods. Experiments on six datasets show our method outperforms the existing ones in identifying all types of multi-view outliers, often by large margins.

NeurIPS Conference 2018 Conference Paper

Representation Learning for Treatment Effect Estimation from Observational Data

  • Liuyi Yao
  • Sheng Li
  • Yaliang Li
  • Mengdi Huai
  • Jing Gao
  • Aidong Zhang

Estimating individual treatment effect (ITE) is a challenging problem in causal inference, due to the missing counterfactuals and the selection bias. Existing ITE estimation methods mainly focus on balancing the distributions of control and treated groups, but ignore the local similarity information that is helpful. In this paper, we propose a local similarity preserved individual treatment effect (SITE) estimation method based on deep representation learning. SITE preserves local similarity and balances data distributions simultaneously, by focusing on several hard samples in each mini-batch. Experimental results on synthetic and three real-world datasets demonstrate the advantages of the proposed SITE method, compared with the state-of-the-art ITE estimation methods.

JBHI Journal 2017 Journal Article

EMG-Torque Relation in Chronic Stroke: A Novel EMG Complexity Representation With a Linear Electrode Array

  • Xu Zhang
  • Dongqing Wang
  • Zaiyang Yu
  • Xiang Chen
  • Sheng Li
  • Ping Zhou

This study examines the electromyogram (EMG)—torque relation for chronic stroke survivors using a novel EMG complexity representation. Ten stroke subjects performed a series of submaximal isometric elbow flexion tasks using their affected and contralateral arms, respectively, while a 20-channel linear electrode array was used to record surface EMG from the biceps brachii muscles. The sample entropy (SampEn) of surface EMG signals was calculated with both global and local tolerance schemes. A regression analysis was performed between SampEn of each channel's surface EMG and elbow flexion torque. It was found that a linear regression can be used to well describe the relation between surface EMG SampEn and the torque. Each channel's root mean square (RMS) amplitude of surface EMG signal in the different torque level was computed to determine the channel with the highest EMG amplitude. The slope of the regression (observed from the channel with the highest EMG amplitude) was smaller on the impaired side than on the nonimpaired side in 8 of the 10 subjects, regardless of the tolerance scheme (global or local) and the range of torques (full or matched range) used for comparison. The surface EMG signals from the channels above the estimated muscle innervation zones demonstrated significantly lower levels of complexity compared with other channels between innervation zones and muscle tendons. The study provides a novel point of view of the EMG-torque relation in the complexity domain, and reveals its alterations post stroke, which are associated with complex neural and muscular changes post stroke. The slope difference between channels with regard to innervation zones also confirms the relevance of electrode position in surface EMG analysis.

IJCAI Conference 2017 Conference Paper

From Ensemble Clustering to Multi-View Clustering

  • Zhiqiang Tao
  • Hongfu Liu
  • Sheng Li
  • Zhengming Ding
  • Yun Fu

Multi-View Clustering (MVC) aims to find the cluster structure shared by multiple views of a particular dataset. Existing MVC methods mainly integrate the raw data from different views, while ignoring the high-level information. Thus, their performance may degrade due to the conflict between heterogeneous features and the noises existing in each individual view. To overcome this problem, we propose a novel Multi-View Ensemble Clustering (MVEC) framework to solve MVC in an Ensemble Clustering (EC) way, which generates Basic Partitions (BPs) for each view individually and seeks for a consensus partition among all the BPs. By this means, we naturally leverage the complementary information of multi-view data in the same partition space. Instead of directly fusing BPs, we employ the low-rank and sparse decomposition to explicitly consider the connection between different views and detect the noises in each view. Moreover, the spectral ensemble clustering task is also involved by our framework with a carefully designed constraint, making MVEC a unified optimization framework to achieve the final consensus partition. Experimental results on six real-world datasets show the efficacy of our approach compared with both MVC and EC methods.

NeurIPS Conference 2017 Conference Paper

Matching on Balanced Nonlinear Representations for Treatment Effects Estimation

  • Sheng Li
  • Yun Fu

Estimating treatment effects from observational data is challenging due to the missing counterfactuals. Matching is an effective strategy to tackle this problem. The widely used matching estimators such as nearest neighbor matching (NNM) pair the treated units with the most similar control units in terms of covariates, and then estimate treatment effects accordingly. However, the existing matching estimators have poor performance when the distributions of control and treatment groups are unbalanced. Moreover, theoretical analysis suggests that the bias of causal effect estimation would increase with the dimension of covariates. In this paper, we aim to address these problems by learning low-dimensional balanced and nonlinear representations (BNR) for observational data. In particular, we convert counterfactual prediction as a classification problem, develop a kernel learning model with domain adaptation constraint, and design a novel matching estimator. The dimension of covariates will be significantly reduced after projecting data to a low-dimensional subspace. Experiments on several synthetic and real-world datasets demonstrate the effectiveness of our approach.

IJCAI Conference 2016 Conference Paper

Learning Robust Representations for Data Analytics

  • Sheng Li

Learning compact representations from high-dimensional and large-scale data plays an essential role in many real-world applications. However, many existing methods show limited performance when data are contaminated with severe noise. To address this challenge, we have proposed several effective methods to extract robust data representations, such as balanced graphs, discriminative subspaces, and robust dictionaries. In addition, several topics are provided as future work.

IJCAI Conference 2016 Conference Paper

Matching via Dimensionality Reduction for Estimation of Treatment Effects in Digital Marketing Campaigns

  • Sheng Li
  • Nikos Vlassis
  • Jaya Kawale
  • Yun Fu

A widely used method for estimating counterfactuals and causal treatment effects from observational data is nearest-neighbor matching. This typically involves pairing each treated unit with its nearest-in-covariates control unit, and then estimating an average treatment effect from the set of matched pairs. Although straightforward to implement, this estimator is known to suffer from a bias that increases with the dimensionality of the covariate space, which can be undesirable in applications that involve high-dimensional data. To address this problem, we propose a novel estimator that first projects the data to a number of random linear subspaces, and it then estimates the median treatment effect by nearest-neighbor matching in each subspace. We empirically compute the mean square error of the proposed estimator using semi-synthetic data, and we demonstrate the method on real-world digital marketing campaign data. The results show marked improvement over baseline methods.

IJCAI Conference 2015 Conference Paper

Cross-View Projective Dictionary Learning for Person Re-Identification

  • Sheng Li
  • Ming Shao
  • Yun Fu

Person re-identification plays an important role in many safety-critical applications. Existing works mainly focus on extracting patch-level features or learning distance metrics. However, the representation power of extracted features might be limited, due to the various viewing conditions of pedestrian images in reality. To improve the representation power of features, we learn discriminative and robust representations via dictionary learning in this paper. First, we propose a cross-view projective dictionary learning (CPDL) approach, which learns effective features for persons across different views. CPDL is a general framework for multiview dictionary learning. Secondly, by utilizing the CPDL framework, we design two objectives to learn low-dimensional representations for each pedestrian in the patch-level and the image-level, respectively. The proposed objectives can capture the intrinsic relationships of different representation coefficients in various settings. We devise efficient optimization algorithms to solve the objectives. Finally, a fusion strategy is utilized to generate the similarity scores. Experiments on the public VIPeR and CUHK Campus datasets show that our approach achieves the state-of-the-art performance.

IJCAI Conference 2015 Conference Paper

Deep Linear Coding for Fast Graph Clustering

  • Ming Shao
  • Sheng Li
  • Zhengming Ding
  • Yun Fu

Clustering has been one of the most critical unsupervised learning techniques that has been widely applied in data mining problems. As one of its branches, graph clustering enjoys its popularity due to its appealing performance and strong theoretical supports. However, the eigen-decomposition problems involved are computationally expensive. In this paper, we propose a deep structure with a linear coder as the building block for fast graph clustering, called Deep Linear Coding (DLC). Different from conventional coding schemes, we jointly learn the feature transform function and discriminative codings, and guarantee that the learned codes are robust in spite of local distortions. In addition, we use the proposed linear coders as the building blocks to formulate a deep structure to further refine features in a layerwise fashion. Extensive experiments on clustering tasks demonstrate that our method performs well in terms of both time complexity and clustering accuracy. On a large-scale benchmark dataset (580K), our method runs 1500 times faster than the original spectral clustering.

YNIMG Journal 2015 Journal Article

Sharpened cortical tuning and enhanced cortico-cortical communication contribute to the long-term neural mechanisms of visual motion perceptual learning

  • Nihong Chen
  • Taiyong Bi
  • Tiangang Zhou
  • Sheng Li
  • Zili Liu
  • Fang Fang

Much has been debated about whether the neural plasticity mediating perceptual learning takes place at the sensory or decision-making stage in the brain. To investigate this, we trained human subjects in a visual motion direction discrimination task. Behavioral performance and BOLD signals were measured before, immediately after, and two weeks after training. Parallel to subjects' long-lasting behavioral improvement, the neural selectivity in V3A and the effective connectivity from V3A to IPS (intraparietal sulcus, a motion decision-making area) exhibited a persistent increase for the trained direction. Moreover, the improvement was well explained by a linear combination of the selectivity and connectivity increases. These findings suggest that the long-term neural mechanisms of motion perceptual learning are implemented by sharpening cortical tuning to trained stimuli at the sensory processing stage, as well as by optimizing the connections between sensory and decision-making areas in the brain.

IJCAI Conference 2013 Conference Paper

Fusion of Word and Letter Based Metrics for Automatic MT Evaluation

  • Muyun Yang
  • Junguo Zhu
  • Sheng Li
  • Tiejun Zhao

With the progress in machine translation, it becomes more subtle to develop the evaluation metric capturing the systems’ differences in comparison to the human translations. In contrast to the current efforts in leveraging more linguistic information to depict translation quality, this paper takes the thread of combining language independent features for a robust solution to MT evaluation metric. To compete with finer granularity of modeling brought by linguistic features, the proposed method augments the word level metrics by a letter based calculation. An empirical study is then conducted over WMT data to train the metrics by ranking SVM. The results reveal that the integration of current language independent metrics can generate well enough performance for a variety of languages. Time-split data validation is promising as a better training setting, though the greedy strategy also works well.

IJCAI Conference 2013 Conference Paper

Low-Rank Coding with b-Matching Constraint for Semi-Supervised Classification

  • Sheng Li
  • Yun Fu

Graph based semi-supervised learning (GSSL) plays an important role in machine learning systems. The most crucial step in GSSL is graph construction. Although several interesting graph construction methods have been proposed in recent years, how to construct an effective graph is still an open problem. In this paper, we develop a novel approach to constructing graph, which is based on low-rank coding and b-matching constraint. By virtue of recent advances in low-rank subspace recovery theory, compact encoding using low-rank representation coefficients allows us to obtain a robust similarity metric between all pairs of samples. Meanwhile, the b-matching constraint helps in obtaining a sparse and balanced graph, which benefits label propagation in GSSL. We build a joint optimization model to learn lowrank codes and balanced graph simultaneously. After using a graph re-weighting strategy, we present a semi-supervised learning algorithm by incorporating our sparse and balanced graph with Gaussian harmonic function (GHF). Experimental results on the Extended YaleB, PIE, ORL and USPS databases demonstrate that our graph outperforms several state-of-the-art graphs, especially when the labeled samples are very scarce.

TIST Journal 2011 Journal Article

Two-Word Collocation Extraction Using Monolingual Word Alignment Method

  • Zhanyi Liu
  • Haifeng Wang
  • Hua Wu
  • Sheng Li

Statistical bilingual word alignment has been well studied in the field of machine translation. This article adapts the bilingual word alignment algorithm into a monolingual scenario to extract collocations from monolingual corpus, based on the fact that the words in a collocation tend to co-occur in similar contexts as in bilingual word alignment. First, the monolingual corpus is replicated to generate a parallel corpus, in which each sentence pair consists of two identical sentences. Next, the monolingual word alignment algorithm is employed to align potentially collocated words. Finally, the aligned word pairs are ranked according to the alignment scores and candidates with higher scores are extracted as collocations. We conducted experiments on Chinese and English corpora respectively. Compared to previous approaches that use association measures to extract collocations from co-occurrence word pairs within a given window, our method achieves higher precision and recall. According to human evaluation, our method achieves precisions of 62% on a Chinese corpus and 64% on an English corpus. In particular, we can extract collocations with longer spans, achieving a higher precision of 83% on the long-span (> 6 words) Chinese collocations.

EAAI Journal 2010 Journal Article

Representation of functional micro-knowledge cell (FMKC) for conceptual design

  • Sheng Li
  • Jie Hu
  • Ying-Hong Peng

The conceptual design process of a complex product concerns multidisciplinary design knowledge. This paper presents an approach of consistent knowledge representation for conceptual design. Firstly, the concept of functional micro-knowledge cell (FMKC) is presented, and the knowledge representation approach of FMKC is proposed. In the function layer, the functional ontology is used to provide a rich vocabulary. By constructing the mapping relationships between the function and structure layer, a systematic knowledge representation scheme is obtained. Secondly, the functional knowledge decomposition theory is proposed to unify the resolution of knowledge representation. Finally, the FMKC is applied to a hydraulic cylinder, which demonstrates the possibility in representing the multidisciplinary knowledge for conceptual design with a unified systematic representation scheme. In addition, the FMKC can be used in knowledge fusion and design reuse for our further study.

IJCAI Conference 2007 Conference Paper

  • Shiqi Zhao
  • Ting Liu
  • Xincheng Yuan
  • Sheng Li
  • Yu Zhang

Lexical paraphrasing aims at acquiring word-level paraphrases. It is critical for many Natural Language Processing (NLP) applications, such as Question Answering (QA), Information Extraction (IE), and Machine Translation (MT). Since the meaning and usage of a word can vary in distinct contexts, different paraphrases should be acquired according to the contexts. However, most of the existing researches focus on constructing paraphrase corpora, in which little contextual constraints for paraphrase application are imposed. This paper presents a method that automatically acquires context-specific lexical paraphrases. In this method, the obtained paraphrases of a word depend on the specific sentence the word occurs in. Two stages are included, i. e. candidate paraphrase extraction and paraphrase validation, both of which are mainly based on web mining. Evaluations are conducted on a news title corpus and the presented method is compared with a paraphrasing method that exploits a Chinese thesaurus of synonyms -- Tongyi Cilin (Extended) (CilinE for short). Results show that the f-measure of our method (0. 4852) is significantly higher than that using CilinE (0. 1127). In addition, over 85% of the correct paraphrases derived by our method cannot be found in CilinE, which suggests that our method is effective in acquiring out-of-thesaurus paraphrases.

v2026.09.13