Arrow Research search

Author name cluster

Jing Lu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

19 papers
2 author rows

Possible papers

19

AAAI Conference 2026 Conference Paper

PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement

  • Xiaobin Rong
  • Qinwen Hu
  • Mansur Yesilbursa
  • Kamil Wojcicki
  • Jing Lu

Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional discriminative approaches. However, existing generative SE approaches often overlook the risk of hallucination under severe noise, leading to incorrect spoken content or inconsistent speaker characteristics, which we term linguistic and acoustic hallucinations, respectively. We argue that linguistic hallucination stems from models' failure to constrain valid phonological structures and it is a more fundamental challenge. While language models (LMs) are well-suited for capturing the underlying speech structure through modeling the distribution of discrete tokens, existing approaches are limited in learning from noise-corrupted representations, which can lead to contaminated priors and hallucinations. To overcome these limitations, we propose the Phonologically Anchored Speech Enhancer (PASE), a generative SE framework that leverages the robust phonological prior embedded in the pre-trained WavLM model to mitigate hallucinations. First, we adapt WavLM into a denoising expert via representation distillation to clean its final-layer features. Guided by the model's intrinsic phonological prior, this process enables robust denoising while minimizing linguistic hallucinations. To further reduce acoustic hallucinations, we train the vocoder with a dual-stream representation: the high-level phonetic representation provides clean linguistic content, while a low-level acoustic representation retains speaker identity and prosody. Experimental results demonstrate that PASE not only surpasses state-of-the-art discriminative models in perceptual quality, but also significantly outperforms prior generative models with substantially lower linguistic and acoustic hallucinations.

AAAI Conference 2026 Conference Paper

Refine3D: Scene-Adaptive Reference Point Refinement for Sparse 3D Object Detection

  • Fan Li
  • Jing Lu
  • Yunlu Xu
  • Changhong Wu
  • Tao Xu
  • Zhaoyi Xiang
  • Yi Niu

Sparse query-based detectors have emerged as the dominant paradigm in camera-only 3D object detection, owing to their exceptional performance and computational efficiency. A central component of these approaches is the use of reference points, which serve as learnable spatial anchors to guide queries in localizing target objects. However, existing methods typically employ a unified set of reference points across all scenes, a design we find suboptimal for handling complex scenarios with highly imbalanced object distributions, such as road intersections or occluded environments. In this paper, we investigate the adaptability of reference points and propose Refine3D, an adaptive refinement mechanism that achieves scene-level alignment between the distribution of reference points and ground-truth objects. In particular, we introduce a novel Reference Point Distribution Loss (RPD-Loss) to ensure reference points converge globally toward object positions, and a Scene-Adaptive Refinement head (SAR-Head) that predicts dynamic offsets for each reference point. Both components can be seamlessly integrated into mainstream sparse detectors. Extensive experiments on two challenging autonomous driving datasets demonstrate that Refine3D outperforms the state-of-the-art with improved detection accuracy and robustness.

AAAI Conference 2026 Conference Paper

Rethinking Flow and Diffusion Bridge Models for Speech Enhancement

  • Dahan Wang
  • Jun Gao
  • Tong Lei
  • Yuxiang Hu
  • Changbao Zhu
  • Kai Chen
  • Jing Lu

Flow matching and diffusion bridge models have emerged as leading paradigms in generative speech enhancement, modeling stochastic processes between paired noisy and clean speech signals based on principles such as flow matching, score matching, and Schrödinger bridge. In this paper, we present a framework that unifies existing flow and diffusion bridge models by interpreting them as constructions of Gaussian probability paths with varying means and variances between paired data. Furthermore, we investigate the underlying consistency between the training/inference procedures of these generative models and conventional predictive models. Our analysis reveals that each sampling step of a well-trained flow or diffusion bridge model optimized with a data prediction loss is theoretically analogous to executing predictive speech enhancement. Motivated by this insight, we introduce an enhanced bridge model that integrates an effective probability path design with key elements from predictive paradigms, including improved network architecture, tailored loss functions, and optimized training strategies. Experiments on denoising and dereverberation tasks demonstrate that the proposed method outperforms existing flow and diffusion baselines with fewer parameters and reduced computational complexity. The results also highlight that the inherently predictive nature of this generative framework imposes limitations on its achievable upper-bound performance.

JBHI Journal 2025 Journal Article

An Eye Video Oriented and rPPG-Based Intraocular Pressure Detection Method

  • Kun Zheng
  • Xuejia Zhen
  • Boxiang Hu
  • Guang Chen
  • Xinming Peng
  • Jing Lu
  • Peng Chen

Current intraocular pressure (IOP) measurement methods still primarily rely on contact IOP measurement instruments, which are inconvenient for widespread use. This study proposed an innovative method for detecting and classifying IOP using eye videos, based on remote photoplethysmography (rPPG). The IOP-Net model was developed by extracting blood volume pulse (BVP) signals from three regions of interest (ROI)—the pupil, iris and sclera, and training a convolutional neural network (CNN) with four convolutional layers. This model can be used to detect the IOP and determine the classification of normal IOP and high IOP. The root mean square errors (RMSE) on EVIP-1 and EVIP-2 datasets were 3. 14 mmHg and 4. 19 mmHg, respectively. When the ground truth of IOP is more than 30 mmHg, the accuracy of the model in classifying high IOP reaches 80. 25%. The results indicate that this method has promising and potential application for video-based IOP detection and classification.

IJCAI Conference 2025 Conference Paper

BridgeVoC: Neural Vocoder with Schrödinger Bridge

  • Tong Lei
  • Zhiyu Zhang
  • Rilin Chen
  • Meng Yu
  • Jing Lu
  • Chengshi Zheng
  • Dong Yu
  • Andong Li

While previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restoration task, this paper proposes a time-frequency (T-F) domain-based neural vocoder with the Schrödinger Bridge, called BridgeVoC, which is the first to follow the data-to-data generation paradigm. Specifically, the mel-spectrogram can be projected into the target linear-scale domain and regarded as a degraded spectral representation with a deficient rank distribution. Based on this, the Schrödinger Bridge is leveraged to establish a connection between the degraded and target data distributions. During the inference stage, starting from the degraded representation, the target spectrum can be gradually restored rather than generated from a Gaussian noise process. Quantitative experiments on LJSpeech and LibriTTS show that BridgeVoC achieves faster inference and surpasses existing diffusion-based vocoder baselines, while also matching or exceeding non-diffusion state-of-the-art methods across evaluation metrics.

AAAI Conference 2025 Conference Paper

Towards Multimodal Sentiment Analysis via Hierarchical Correlation Modeling with Semantic Distribution Constraints

  • Qinfu Xu
  • Yiwei Wei
  • Chunlei Wu
  • Leiquan Wang
  • Shaozu Yuan
  • Jie Wu
  • Jing Lu
  • Hengyang Zhou

Sentiment analysis is rapidly advancing by utilizing various data modalities (e.g., text, video, and audio). However, most existing techniques only learn the atomic-level features that reflect strong correlations, while ignoring more complex compositions in multimodal data. Moreover, they also neglected the incongruity in semantic distribution among modalities. In light of this, we introduce a novel Hierarchical Correlation Modeling Network (HCMNet), which enhances the multimodal sentiment analysis by exploring both the atomic-level correlations based on dynamic attention reasoning and the composition-level correlations through topological graph reasoning. In addition, we also alleviate the impact of distributional inconsistencies between modalities from both atomic-level and composition-level perspectives. Specifically, we first design an atomic-level contrastive loss that constrains the semantic distribution across modalities to mitigate the atomic-level inconsistency. Then, we design a graph optimal transport module that integrates transport flows with different graphs to constrain the composition-level semantic distribution, thus reducing the inconsistency of compositional nodes. Experiments on three public benchmark datasets have demonstrated the superiority of the proposed model over the state-of-the-art methods.

EAAI Journal 2024 Journal Article

Multi-receptive Field Distillation Network for seismic velocity model building

  • Jing Lu
  • Chunlei Wu
  • Jianping Huang
  • Guolong Li
  • Shaozu Yuan

Velocity model building is crucial for seismic exploration, yet conventional methods struggle with complex geological scenarios due to assumptions of horizontal layering. These challenges are exacerbated in areas with complex structures and low signal-to-noise ratios, where precise velocity determination is difficult. To overcome these limitations, we propose the Multi-receptive Field Distillation Network, a novel deep-learning approach that leverages Multi-receptive Field modules to extract detailed seismic features, and a Shot Record Transformer Block to capture long-range dependencies. Our network employs a distillation architecture to process seismic data with and without noise, enhancing the model’s ability to learn from noisy records. A Composite Loss function is introduced for optimizing model parameters, promoting a unified feature representation. Numerical experiments on synthetic models demonstrate our method’s superior performance in noise inversion and velocity modeling accuracy. The results underscore the network’s potential for improving seismic exploration accuracy. The code has been made public on GitHub: UPCvmb/MFD. git

NeurIPS Conference 2023 Conference Paper

Learning List-Level Domain-Invariant Representations for Ranking

  • Ruicheng Xian
  • Honglei Zhuang
  • Zhen Qin
  • Hamed Zamani
  • Jing Lu
  • Ji Ma
  • Kai Hui
  • Han Zhao

Domain adaptation aims to transfer the knowledge learned on (data-rich) source domains to (low-resource) target domains, and a popular method is invariant representation learning, which matches and aligns the data distributions on the feature space. Although this method is studied extensively and applied on classification and regression problems, its adoption on ranking problems is sporadic, and the few existing implementations lack theoretical justifications. This paper revisits invariant representation learning for ranking. Upon reviewing prior work, we found that they implement what we call item-level alignment, which aligns the distributions of the items being ranked from all lists in aggregate but ignores their list structure. However, the list structure should be leveraged, because it is intrinsic to ranking problems where the data and the metrics are defined and computed on lists, not the items by themselves. To close this discrepancy, we propose list-level alignment—learning domain-invariant representations at the higher level of lists. The benefits are twofold: it leads to the first domain adaptation generalization bound for ranking, in turn providing theoretical support for the proposed method, and it achieves better empirical transfer performance for unsupervised domain adaptation on ranking tasks, including passage reranking.

ICLR Conference 2023 Conference Paper

Promptagator: Few-shot Dense Retrieval From 8 Examples

  • Zhuyun Dai
  • Vincent Y. Zhao
  • Ji Ma 0004
  • Yi Luan
  • Jianmo Ni
  • Jing Lu
  • Anton Bakalov
  • Kelvin Guu

Much recent research on information retrieval has focused on how to transfer from one task (typically with abundant supervised data) to various other retrieval tasks where supervision is limited, with the implicit assumption that it is possible to generalize from one task to all the rest. However, this overlooks the fact that there are many diverse and unique retrieval problems, each targeting different search intents, queries, and search domains. In this paper, we suggest to work on Few-shot Dense Retrieval, a setting where each task comes with a short description and a few examples. To address this, we introduce Prompt-based Query Generation forRetrieval (Promptagator): for each task, we feed the few-shot examples to a large language model (LLM) and prompt it to behave as a task-specific query generator. Using this, we can synthetically generate a large number of relevant queries for any document, yielding abundant data for training task-specific retrievers --- with no reliance on traditional resources such as Natural Questions (Kwiatkowskiet al., 2019) or MS MARCO (Nguyen et al., 2016). Surprisingly, Promptagator with only 8 annotated examples enables efficient dual encoder retrievers to outperform computationally more expensive models trained on MS MARCO such as ColBERT v2 (Santhanam et al., 2022) by more than 1.2 points nDCG@10 on average on 11 retrieval sets. Further training standard-size re-rankers using the same generated data yields another 5.0 points nDCG@10 improvement. Our studies show that synthetic query generation can be far more effective than previously observed, especially when a small amount of task-specific knowledge is given.

YNIMG Journal 2023 Journal Article

Temporal trajectories of normal myelination and axonal development assessed by quantitative macromolecular and diffusion MRI: Ultrastructural and immunochemical validation in a rabbit model

  • Alexander Drobyshevsky
  • Sylvia Synowiec
  • Ivan Goussakov
  • Jing Lu
  • David Gascoigne
  • Daniil P Aksenov
  • Vasily Yarnykh

INTRODUCTION: Quantitative and non-invasive measures of brain myelination and maturation during development are of great importance to both clinical and translational research communities. While the metrics derived from diffusion tensor imaging, are sensitive to developmental changes and some pathologies, they remain difficult to relate to the actual microstructure of the brain tissue. The advent of advanced model-based microstructural metrics requires histological validation. The purpose of the study was to validate novel, model-based MRI techniques, such as macromolecular proton fraction mapping (MPF) and neurite orientation and dispersion indexing (NODDI), against histologically derived indexes of myelination and microstructural maturation at various stages of development. METHODS: New Zealand White rabbit kits underwent serial in-vivo MRI examination at postnatal days 1, 5, 11, 18, and 25, and as adults. Multi-shell, diffusion-weighted experiments were processed to fit NODDI model to obtain estimates, intracellular volume fraction (ICVF) and orientation dispersion index (ODI). Macromolecular proton fraction (MPF) maps were obtained from three source (MT-, PD-, and T1-weighted) images. After MRI sessions, a subset of animals was euthanized and regional samples of gray and white matter were taken for western blot analysis, to determine myelin basic protein (MBP), and electron microscopy, to estimate axonal, myelin fractions and g-ratio. RESULTS: MPF of white matter regions showed a period of fast growth between P5 and P11 in the internal capsule, with a later onset in the corpus callosum. This MPF trajectory was in agreement with levels of myelination in the corresponding brain region, as assessed by western blot and electron microscopy. In the cortex, the greatest increase of MPF occurred between P18 and P26. In contrast, myelin, according to MBP western blot, saw the largest hike between P5 and P11 in the sensorimotor cortex and between P11 and P18 in the frontal cortex, which then seemingly plateaued after P11 and P18 respectively. G-ratio by MRI markers decreased with age in the white matter. However, electron microscopy suggest a relatively stable g-ratio throughout development. CONCLUSION: Developmental trajectories of MPF accurately reflected regional differences of myelination rate in different cortical regions and white matter tracts. MRI-derived estimation of g-ratio was inaccurate during early development, likely due to the overestimation of axonal volume fraction by NODDI due to the presence of a large proportion of unmyelinated axons.

AAAI Conference 2022 Conference Paper

PMAL: Open Set Recognition via Robust Prototype Mining

  • Jing Lu
  • Yunlu Xu
  • Hao Li
  • Zhanzhan Cheng
  • Yi Niu

Open Set Recognition (OSR) has been an emerging topic. Besides recognizing predefined classes, the system needs to reject the unknowns. Prototype learning is a potential manner to handle the problem, as its ability to improve intra-class compactness of representations is much needed in discrimination between the known and the unknowns. In this work, we propose a novel Prototype Mining And Learning (PMAL) framework. It has a prototype mining mechanism before the phase of optimizing embedding space, explicitly considering two crucial properties, namely high-quality and diversity of the prototype set. Concretely, a set of high-quality candidates are firstly extracted from training samples based on data uncertainty learning, avoiding the interference from unexpected noise. Considering the multifarious appearance of objects even in a single category, a diversity-based strategy for prototype set filtering is proposed. Accordingly, the embedding space can be better optimized to discriminate therein the predefined classes and between known and unknowns. Extensive experiments verify the two good characteristics (i. e. , high-quality and diversity) embraced in prototype mining, and show the remarkable performance of the proposed framework compared to state-of-the-arts.

TCS Journal 2022 Journal Article

The monad on strong quasi-metric spaces

  • Jing Lu

In this paper, we introduce the notion of strong quasi-metric spaces, and prove that the formal ball construction B induces monads on the category of strong quasi-metric spaces with 1-Lipschitz maps as morphisms, and on the category of strong quasi-metric spaces with Y-continuous maps as morphisms, respectively.

AAAI Conference 2021 Conference Paper

Span-Based Event Coreference Resolution

  • Jing Lu
  • Vincent Ng

Motivated by the recent successful application of span-based models to entity-based information extraction tasks, we investigate span-based models for event coreference resolution, focusing on determining (1) whether the successes of spanbased models of entity coreference can be extended to event coreference; (2) whether exploiting the dependency between event coreference and the related subtask of trigger detection; and (3) whether automatically computed entity coreference information can benefit span-based event coreference resolution. Empirical results on the standard evaluation dataset provide affirmative answers to all three questions.

NeurIPS Conference 2020 Conference Paper

Kalman Filtering Attention for User Behavior Modeling in CTR Prediction

  • Hu Liu
  • Jing Lu
  • Xiwei Zhao
  • Sulong Xu
  • Hao Peng
  • Yutong Liu
  • Zehua Zhang
  • Jian Li

Click-through rate (CTR) prediction is one of the fundamental tasks for e-commerce search engines. As search becomes more personalized, it is necessary to capture the user interest from rich behavior data. Existing user behavior modeling algorithms develop different attention mechanisms to emphasize query-relevant behaviors and suppress irrelevant ones. Despite being extensively studied, these attentions still suffer from two limitations. First, conventional attentions mostly limit the attention field only to a single user's behaviors, which is not suitable in e-commerce where users often hunt for new demands that are irrelevant to any historical behaviors. Second, these attentions are usually biased towards frequent behaviors, which is unreasonable since high frequency does not necessarily indicate great importance. To tackle the two limitations, we propose a novel attention mechanism, termed Kalman Filtering Attention (KFAtt), that considers the weighted pooling in attention as a maximum a posteriori (MAP) estimation. By incorporating a priori, KFAtt resorts to global statistics when few user behaviors are relevant. Moreover, a frequency capping mechanism is incorporated to correct the bias towards frequent behaviors. Offline experiments on both benchmark and a 10 billion scale real production dataset, together with an Online A/B test, show that KFAtt outperforms all compared state-of-the-arts. KFAtt has been deployed in the ranking system of JD. com, one of the largest B2C e-commerce websites in China, serving the main traffic of hundreds of millions of active users.

IJCAI Conference 2019 Conference Paper

Position Focused Attention Network for Image-Text Matching

  • Yaxiong Wang
  • Hao Yang
  • Xueming Qian
  • Lin Ma
  • Jing Lu
  • Biao Li
  • Xin Fan

Image-text matching tasks have recently attracted a lot of attention in the computer vision field. The key point of this cross-domain problem is how to accurately measure the similarity between the visual and the textual contents, which demands a fine understanding of both modalities. In this paper, we propose a novel position focused attention network (PFAN) to investigate the relation between the visual and the textual views. In this work, we integrate the object position clue to enhance the visual-text joint-embedding learning. We first split the images into blocks, by which we infer the relative position of region in the image. Then, an attention mechanism is proposed to model the relations between the image region and blocks and generate the valuable position feature, which will be further utilized to enhance the region expression and model a more reliable relationship between the visual image and the textual sentence. Experiments on the popular datasets Flickr30K and MS-COCO show the effectiveness of the proposed method. Besides the public datasets, we also conduct experiments on our collected practical news dataset (Tencent-News) to validate the practical application value of proposed method. As far as we know, this is the first attempt to test the performance on the practical application. Our method can achieve the state-of-art performance on all of these three datasets.

IJCAI Conference 2018 Conference Paper

Event Coreference Resolution: A Survey of Two Decades of Research

  • Jing Lu
  • Vincent Ng

Recent years have seen a gradual shift of focus from entity-based tasks to event-based tasks in information extraction research. Being a core event-based task, event coreference resolution is less studied but arguably more challenging than entity coreference resolution. This paper provides an overview of the major milestones made in event coreference research since its inception two decades ago.

IJCAI Conference 2018 Conference Paper

Online Deep Learning: Learning Deep Neural Networks on the Fly

  • Doyen Sahoo
  • Quang Pham
  • Jing Lu
  • Steven C. H. Hoi

Deep Neural Networks (DNNs) are typically trained by backpropagation in a batch setting, requiring the entire training data to be made available prior to the learning task. This is not scalable for many real-world scenarios where new data arrives sequentially in a stream. We aim to address an open challenge of ``Online Deep Learning" (ODL) for learning DNNs on the fly in an online setting. Unlike traditional online learning that often optimizes some convex objective function with respect to a shallow model (e. g. , a linear/kernel-based hypothesis), ODL is more challenging as the optimization objective is non-convex, and regular DNN with standard backpropagation does not work well in practice for online settings. We present a new ODL framework that attempts to tackle the challenges by learning DNN models which dynamically adapt depth from a sequence of training data in an online learning setting. Specifically, we propose a novel Hedge Backpropagation (HBP) method for online updating the parameters of DNN effectively, and validate the efficacy on large data sets (both stationary and concept drifting scenarios).

TIST Journal 2018 Journal Article

Sparse Passive-Aggressive Learning for Bounded Online Kernel Methods

  • Jing Lu
  • Doyen Sahoo
  • Peilin Zhao
  • Steven C. H. Hoi

One critical deficiency of traditional online kernel learning methods is their unbounded and growing number of support vectors in the online learning process, making them inefficient and non-scalable for large-scale applications. Recent studies on scalable online kernel learning have attempted to overcome this shortcoming, e.g., by imposing a constant budget on the number of support vectors. Although they attempt to bound the number of support vectors at each online learning iteration, most of them fail to bound the number of support vectors for the final output hypothesis, which is often obtained by averaging the series of hypotheses over all the iterations. In this article, we propose a novel framework for bounded online kernel methods, named “Sparse Passive-Aggressive (SPA)” learning, which is able to yield a final output kernel-based hypothesis with a bounded number of support vectors. Unlike the common budget maintenance strategy used by many existing budget online kernel learning approaches, the idea of our approach is to attain the bounded number of support vectors using an efficient stochastic sampling strategy that samples an incoming training example as a new support vector with a probability proportional to its loss suffered. We theoretically prove that SPA achieves an optimal mistake bound in expectation, and we empirically show that it outperforms various budget online kernel learning algorithms. Finally, in addition to general online kernel learning tasks, we also apply SPA to derive bounded online multiple-kernel learning algorithms, which can significantly improve the scalability of traditional Online Multiple-Kernel Classification (OMKC) algorithms while achieving satisfactory learning accuracy as compared with the existing unbounded OMKC algorithms.

JMLR Journal 2016 Journal Article

Large Scale Online Kernel Learning

  • Jing Lu
  • Steven C.H. Hoi
  • Jialei Wang
  • Peilin Zhao
  • Zhi-Yong Liu

In this paper, we present a new framework for large scale online kernel learning, making kernel methods efficient and scalable for large-scale online learning applications. Unlike the regular budget online kernel learning scheme that usually uses some budget maintenance strategies to bound the number of support vectors, our framework explores a completely different approach of kernel functional approximation techniques to make the subsequent online learning task efficient and scalable. Specifically, we present two different online kernel machine learning algorithms: (i) Fourier Online Gradient Descent (FOGD) algorithm that applies the random Fourier features for approximating kernel functions; and (ii) Nyström Online Gradient Descent (NOGD) algorithm that applies the Nyström method to approximate large kernel matrices. We explore these two approaches to tackle three online learning tasks: binary classification, multi-class classification, and regression. The encouraging results of our experiments on large-scale datasets validate the effectiveness and efficiency of the proposed algorithms, making them potentially more practical than the family of existing budget online kernel learning approaches. [abs] [ pdf ][ bib ] &copy JMLR 2016. ( edit, beta )

v2026.09.13