Arrow Research search

Author name cluster

Fang Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

37 papers
1 author row

Possible papers

37

AAAI Conference 2026 Conference Paper

Evolving Semantic Propagation for Aerial Semantic 3D Gaussian Splatting

  • Zihan Gao
  • Lingling Li
  • Xu Liu
  • Fang Liu
  • Licheng Jiao
  • Puhua Chen
  • Wenping Ma
  • Shuyuan Yang

Semantic understanding of large-scale aerial scenes represents a critical challenge in 3D computer vision, hindered by the prohibitive cost of dense annotation. This paper introduces EvoPropGS, a novel approach for the semantic segmentation of 3D Gaussian Splatting models that requires only minimal supervision. Our core insight is to leverage the inherent structural repetitions within aerial environments to propagate semantic information from a sparse set of annotations across the entire 3D scene. Our approach constructs a prompt library by pairing SAM-generated mask candidates with DINOv2 feature embeddings from annotated views. For unannotated regions, we generate pseudo-labels by matching region proposals with these featured prompts via cosine similarity. We then formulate optimal prompt selection as a discrete optimization problem solved via evolutionary search, guided by our novel fitness function that evaluates both 3D consistency and 2D semantic coherence. Extensive experiments demonstrate that EvoPropGS achieves accurate segmentation with only 2 percent annotated pixels.

YNIMG Journal 2026 Journal Article

Functional gradient alteration and structural remodeling in postpartum women

  • Shiyu Xia
  • Xinyu Zhao
  • Bin Lv
  • Yuanyuan Gan
  • Yukun Kang
  • Jiang Long
  • Fang Liu
  • Xiao Hu

Postpartum women (PW) undergo profound brain functional and structural reorganization to support maternal adaptation. However, the specific large-scale neural adaptation mechanisms remain unclear. The current study employed a multimodal MRI approach integrating functional gradient analysis, graph-theoretical network metrics, and morphometry to explore the brain connectome reorganization across the postpartum period and its clinical correlates in 209 participants (134 PW and 75 healthy nulliparous women (HNW)). Compared to HNW, PW exhibited a significant contraction of the first two principal functional gradients, reduced local network segregation and less efficient information processing, accompanied by gray matter volume (GMV) reductions. Mediation analysis revealed that GMV alterations in PW modulate functional gradient reorganization by influencing network integration and segregation. These neural changes were closely linked to clinical symptoms including sleep quality and anxiety. Our findings revealed a large-scale network reconfiguration in PW, simultaneously elucidating neurobiological mechanisms of adaptive plasticity in postpartum period.

AAAI Conference 2026 Conference Paper

HTTrack: Learning to Perceive Targets via Historical Trajectories in Satellite Video Tracking

  • Jiahao Wang
  • Fang Liu
  • Licheng Jiao
  • Hao Wang
  • Shuo Li
  • Xinyi Wang
  • Lingling Li
  • Puhua Chen

In recent years, the rapid progress of deep learning has driven notable advancements in satellite video tracking, a critical task for applications such as environmental monitoring, disaster management, and defense. Despite these strides, existing approaches remain constrained by their inability to handle dynamic challenges, such as target appearance variations, complex motion patterns, and occlusions. Traditional methods often suffer from static template matching or overly complex update mechanisms, compromising their robustness and practicality in real-world scenarios. To address these limitations, we propose a paradigm shift in satellite video tracking by integrating historical trajectory knowledge with visual features. This fusion enhances the tracker's perceptual understanding of targets over time, enabling more adaptive and resilient tracking. By aligning spatial, temporal, and cross-modal information, our approach effectively bridges the gap between fragmented observations and coherent tracking performance, even under challenging conditions like small target detection and cluttered backgrounds. Extensive experiments conducted on multiple satellite video tracking benchmarks demonstrate the superiority of our method, with HTTrack achieving success rates of 51.5% on SV248S, 52.9% on SatSOT, and 32.6% on VISO, significantly outperforming state-of-the-art trackers and marking a step forward in achieving robust, accurate, and scalable satellite video tracking.

AAAI Conference 2026 Conference Paper

LandCraft: Designing the Structured 3D Landscapes via Text Guidance

  • Zhihao Liu
  • Fang Liu
  • Weihao Xuan
  • Naoto YOKOYA

Modeling large-scale landscapes is a foundational yet time-consuming task in many 3D applications, typically requiring substantial expertise. Recently, Text-to-3D techniques have emerged as a promising, beginner-friendly prototyping approach for generating 3D content from textual input. However, existing methods either produce unusable, problematic geometries, or fail to fully capture the user's complex intent from the input text—making it difficult to generate high-quality landscape assets with controllable spatial and geographic features. In this paper, we present LandCraft, a novel AI-assisted authoring tool that enables the rapid creation of high-quality landscape scenes based on user descriptions. Our system employs a coarse-to-fine generation process: Initially, large language and deep generative models concretize textual ideas into abstract representations that capture essential landscape features, such as spatial and geographic characteristics. Then, we leverage a comprehensive procedural generation module to synthesize the detailed, structurally consistent 3D landscapes based on these inferred representations. LandCraft can effectively generate production-ready 3D scene assets that can be seamlessly exported to external game engines or modeling software, enabling immediate practical use.

AAAI Conference 2026 Conference Paper

Preference Optimization via Contrastive Divergence: Your Policy Is Secretly an NLL Estimator

  • Zhuotong Chen
  • Fang Liu
  • Xuan Zhu
  • Haozhu Wang
  • Jiayu Li
  • Yanjun Qi
  • Mohammad Ghavamzadeh

Existing studies on preference optimization (PO) have been focused on constructing pairwise preference data following simple heuristics, such as maximizing the margin between chosen and rejected responses based on human (or AI) ratings. In this work, we develop a novel PO framework that provides theoretical guidance to effectively sample rejected responses. To achieve this, we formulate PO as minimizing the negative log-likelihood (NLL) of a probability model and propose a sampling-based solution to estimate its normalization constant via contrastive divergence. We show that these estimative samples can act as rejected responses in PO. Leveraging the connection established between PO and NLL estimation, we propose a novel PO algorithm, called Monte-Carlo-based PO (MC-PO), that applies a MC kernel to sample *hard negatives* w.r.t.~the log-likelihood of the target policy. Intuitively, these hard negatives represent the rejected samples that are most difficult for the current policy to differentiate. We show that MC-PO outperforms existing SOTA baselines on popular alignment benchmarks.

AAAI Conference 2026 Conference Paper

Semantic Feature Purification for Adversarially-Aware RGB-T Tracking

  • Jiahao Wang
  • Fang Liu
  • Hao Wang
  • Shuo Li
  • Xinyi Wang
  • Puhua Chen

RGB-T tracking is increasingly deployed in safety-critical applications such as autonomous driving, surveillance, and rescue robotics, where tracking reliability is essential under adverse conditions. Although the fusion of RGB and thermal infrared (TIR) modalities offers improved robustness in low-light and occluded scenes, recent findings show that RGB-T trackers remain highly susceptible to subtle input perturbations, human-imperceptible modifications that exploit cross-modal inconsistencies to mislead tracking outputs. In real-world scenarios, such perturbations can arise from sensor spoofing, infrared camouflage, or physical-world attacks, posing serious risks to operational safety. To address this, we propose SFPT, a Semantic Feature Purification framework that enhances RGB-T tracking at the representation level. Rather than filtering corrupted inputs at the pixel level, SFPT introduces task-specific semantic anchors into the feature space to reinforce perturbation-invariant cues. These anchors are derived from descriptive language, interact with visual features to purify representations. To further suppress modality-specific interference, we design an Adaptive Perturbation-Guided Cross-Modal Fusion (APG-CMF) module, which leverages language and visual signals to estimate reliability and dynamically reweight cross-modal features, ensuring robust fusion under perturbation conditions. Extensive experiments under diverse perturbation conditions validate the effectiveness of our approach. Notably, SFPT maintains performance comparable to clean settings even when subjected to perturbations of strength 1/255 and 4/255, demonstrating strong resilience to real-world interference.

YNIMG Journal 2026 Journal Article

When More Control Means Better Choices: Cognitive Control Networks Drive Expected-Value Maximization Under Uncertainty

  • Xia Wu
  • Yuning Geng
  • Yan Chen
  • Shuoxian Zhang
  • Tianhao Liu
  • Shuaipeng You
  • Fang Liu
  • Yunpeng Jiang

Human decision-making under outcome uncertainty often deviates from rational expected-value maximization, frequently falling back on the suboptimal probability matching heuristic. The neurocomputational mechanisms determining individual differences in overcoming this heuristic remain elusive. Here, we investigated how cognitive control capacity (CCC) modulates decision-making under varying levels of outcome uncertainty. Participants with high and low CCC performed a predictive inference task during functional magnetic resonance imaging. Behaviorally, high CCC individuals consistently exhibited a significantly higher proportion of maximizing responses (PMR) across all uncertainty levels. Using hierarchical drift-diffusion modeling, we demonstrated that this optimal performance was driven by more cautious decision thresholds, indicating greater deliberation to resist intuitive shortcuts. At the neural level, while localized activations in the cingulo-opercular network (CON) and frontoparietal network (FPN) reflected the general cognitive burden of escalating uncertainty, functional connectivity analyses revealed a specific neural pathway supporting optimal choices. Crucially, the connectivity within the CON (anterior insula to middle frontal gyrus) acted as a specific neural amplifier, which was absolutely necessary for translating high cognitive capacity into optimal expected-value maximization. The FPN, while tracking uncertainty, did not modulate this capacity-performance link. Together, these findings provide an integrated neurocomputational framework demonstrating how specific cognitive control networks mobilize resources to overcome heuristic tendencies and achieve optimal decisions under uncertainty.

AAAI Conference 2025 Conference Paper

ALLVB: All-in-One Long Video Understanding Benchmark

  • Xichen Tan
  • Yuanjing Luo
  • Yunfan Ye
  • Fang Liu
  • Zhiping Cai

From image to video understanding, the capabilities of Multi-modal LLMs (MLLMs) are increasingly powerful. However, most existing video understanding benchmarks are relatively short, which makes them inadequate for effectively evaluating the long-sequence modeling capabilities of MLLMs. This highlights the urgent need for a comprehensive and integrated long video understanding benchmark to assess the ability of MLLMs thoroughly. To this end, we propose ALLVB (ALL-in-One Long Video Understanding Benchmark). ALLVB's main contributions include: 1) It integrates 9 major video understanding tasks. These tasks are converted into video QA formats, allowing a single benchmark to evaluate 9 different video understanding capabilities of MLLMs, highlighting the versatility, comprehensiveness, and challenging nature of ALLVB. 2) A fully automated annotation pipeline using GPT-4o is designed, requiring only human quality control, which facilitates the maintenance and expansion of the benchmark. 3) It contains 1,376 videos across 16 categories, averaging nearly 2 hours each, with a total of 252k QAs. To the best of our knowledge, it is the largest long video understanding benchmark in terms of the number of videos, average duration, and number of QAs. We have tested various mainstream MLLMs on ALLVB, and the results indicate that even the most advanced commercial models have significant room for improvement. This reflects the benchmark's challenging nature and demonstrates the substantial potential for development in long video understanding.

IJCAI Conference 2025 Conference Paper

fairGNN-WOD: Fair Graph Learning Without Complete Demographics

  • Zichong Wang
  • Fang Liu
  • Shimei Pan
  • Jun Liu
  • Fahad Saeed
  • Meikang Qiu
  • Wenbin Zhang

Graph Neural Networks (GNNs) have excelled in diverse applications due to their outstanding predictive performance, yet they often overlook fairness considerations, prompting numerous recent efforts to address this societal concern. However, most fair GNNs assume complete demographics by design, which is impractical in most real-world socially sensitive applications due to privacy, legal, or regulatory restrictions. For example, the Consumer Financial Protection Bureau (CFPB) mandates that creditors ensure fairness without requesting or collecting information about an applicant’s race, religion, nationality, sex, or other demographics. To this end, this paper proposes fairGNN-WOD, a first-of-its-kind framework that considers mitigating unfairness in graph learning without using demographic information. In addition, this paper provides a theoretical perspective on analyzing bias in node representations and establishes the relationship between utility and fairness objectives. Experiments on three real-world graph datasets illustrate that fairGNN-WOD outperforms state-of-the-art baselines in achieving fairness but also maintains comparable prediction performance.

IJCAI Conference 2025 Conference Paper

Language-Guided Hybrid Representation Learning for Visual Grounding on Remote Sensing Images

  • Biao Liu
  • Xu Liu
  • Lingling Li
  • Licheng Jiao
  • Fang Liu
  • Xinyu Sun
  • Youlin Huang

Visual grounding (VG) refers to detecting the specific objects in images based on linguistic expressions, and it has profound significance in the advanced interpretation of natural images. In remote sensing image interpretation, visual grounding is limited by characteristics such as the complex scenes and diverse object sizes. To solve this problem, we propose a novel remote sensing visual grounding (RSVG) framework, named language-guided hybrid representation learning Transformer (LGFormer). Specifically, we designed a multimodal dual-encoder Transformer structure called the adaptive multimodal feature fusion module. This structure innovatively integrates text and visual features as hybrid queries, enabling early-stage decoding queries to perceive the target position accurately. Then, the different modal information from the dual encoders is aggregated by hybrid queries to obtain the final object embedding for coordinate regression. Besides, a multi-scale cross-modal feature enhancement module (MSCM) is designed to enhance the self-representation of the extracted text and visual features and align them semantically. As for the hybrid queries, we use linguistic guidance to select visual features as the visual part and sentence-level features as the textual part. Finally, the LGFormer model we designed achieved the best results compared to existing models on the DIOR-RSVG and OPT-RSVG datasets.

EAAI Journal 2025 Journal Article

Physics descriptors enhanced Bayesian learning method for permeability of random media under sparse data

  • Hang Qi
  • Xiaofei Guan
  • Qing Chen
  • Zhengwu Jiang
  • Fang Liu
  • Jieqiong Zhang
  • Hehua Zhu

Permeability is a significant property in microstructure-based material design. Currently, the main research gap in such material design can be divided into three aspects: Firstly, experimental methods face stringent conditions and poor reproducibility in accurately describing internal media structures. Secondly, machine learning methods require impractically large datasets to learn the strong nonlinear relationship between permeability and microstructure. Thirdly, there is a challenge in quantifying the inherent uncertainty of randomly distributed phases and the uncertainty propagation in the model. In this work, we propose a novel physics descriptors enhanced Bayesian learning method, aiming to predict the permeability of random media under sparse data. The key aspects of the method are the physics descriptors with clear physical indications, which characterize porous microstructure from multiple perspectives, and an encoder–decoder strategy that reduces microstructural complexity while utilizing these physics descriptors as basis functions to reconstruct permeability. The numerical experiments indicate that porosity, chord length density, and lineal-path dominate in determining permeability. Moreover, predictions with a 5%–11% error can be achieved under 32 training samples. Finally, the method provides a complete posterior distribution, where the posterior variances quantify uncertainty and benefit robust decisions in simulation-assisted materials design.

AAAI Conference 2024 Conference Paper

FG-EmoTalk: Talking Head Video Generation with Fine-Grained Controllable Facial Expressions

  • Zhaoxu Sun
  • Yuze Xuan
  • Fang Liu
  • Yang Xiang

Although deep generative models have greatly improved one-shot video-driven talking head generation, few studies address fine-grained controllable facial expression editing, which is crucial for practical applications. Existing methods rely on a fixed set of predefined discrete emotion labels or simply copy expressions from input videos. This is limiting as expressions are complex, and methods using only emotion labels cannot generate fine-grained, accurate or mixed expressions. Generating talking head video with precise expressions is also difficult using 3D model-based approaches, as 3DMM only models facial movements and tends to produce deviations. In this paper, we propose a novel framework enabling fine-grained facial expression editing in talking face generation. Our goal is to achieve expression control by manipulating the intensities of individual facial Action Units (AUs) or groups. First, compared with existing methods which decouple the face into pose and expression, we propose a disentanglement scheme to isolates three components from the human face, namely, appearance, pose, and expression. Second, we propose to use input AUs to control muscle group intensities in the generated face, and integrate the AUs features with the disentangled expression latent code. Finally, we present a self-supervised training strategy with well-designed constraints. Experiments show our method achieves fine-grained expression control, produces high-quality talking head videos and outperforms baseline methods.

YNIMG Journal 2024 Journal Article

Mothers and fathers show different neural synchrony with their children during shared experiences

  • Qi Liu
  • Siyu Zhu
  • Xinqi Zhou
  • Fang Liu
  • Benjamin Becker
  • Keith M. Kendrick
  • Weihua Zhao

Parent-child shared experiences has an important influence on social development in children although contributions of mothers and fathers may differ. Neural synchronicity occurs between mothers and fathers and their children during social interactions but it is unclear whether they differ in this respect. We used data from simultaneous fNIRS hyperscanning in mothers (n = 33) and fathers (n = 29) and their children (3-4 years) to determine different patterns and strengths of neural synchronization in the frontal cortex during co-viewing of videos or free-play. Mothers showed greater synchrony with child than fathers during passive viewing of videos and the synchronization was positively associated with video complexity and negatively associated with parental stress. During play interactions, mothers showed more controlling behaviors over their child and greater evidence for joint gaze and joint imitation play with child whereas fathers spent more time gazing at other things. In addition, different aspects of child communication promoted neural synchrony between mothers and fathers and child during active play interactions. Overall, our findings indicate greater neural and behavioral synchrony between mothers than fathers and young children during passive or active shared experiences, although for both it was weakened by parental distress and child difficulty.

AAAI Conference 2024 Conference Paper

Multi-View Dynamic Reflection Prior for Video Glass Surface Detection

  • Fang Liu
  • Yuhao Liu
  • Jiaying Lin
  • Ke Xu
  • Rynson W.H. Lau

Recent research has shown significant interest in image-based glass surface detection (GSD). However, detecting glass surfaces in dynamic scenes remains largely unexplored due to the lack of a high-quality dataset and an effective video glass surface detection (VGSD) method. In this paper, we propose the first VGSD approach. Our key observation is that reflections frequently appear on glass surfaces, but they change dynamically as the camera moves. Based on this observation, we propose to offset the excessive dependence on a single uncertainty reflection via joint modeling of temporal and spatial reflection cues. To this end, we propose the VGSD-Net with two novel modules: a Location-aware Reflection Extraction (LRE) module and a Context-enhanced Reflection Integration (CRI) module, for the position-aware reflection feature extraction and the spatial-temporal reflection cues integration, respectively. We have also created the first large-scale video glass surface dataset (VGSD-D), consisting of 19,166 image frames with accurately-annotated glass masks extracted from 297 videos. Extensive experiments demonstrate that VGSD-Net outperforms state-of-the-art approaches adapted from related fields. Code and dataset will be available at https://github.com/fawnliu/VGSD.

AAAI Conference 2024 Conference Paper

Recasting Regional Lighting for Shadow Removal

  • Yuhao Liu
  • Zhanghan Ke
  • Ke Xu
  • Fang Liu
  • Zhenwei Wang
  • Rynson W.H. Lau

Removing shadows requires an understanding of both lighting conditions and object textures in a scene. Existing methods typically learn pixel-level color mappings between shadow and non-shadow images, in which the joint modeling of lighting and object textures is implicit and inadequate. We observe that in a shadow region, the degradation degree of object textures depends on the local illumination, while simply enhancing the local illumination cannot fully recover the attenuated textures. Based on this observation, we propose to condition the restoration of attenuated textures on the corrected local lighting in the shadow region. Specifically, We first design a shadow-aware decomposition network to estimate the illumination and reflectance layers of shadow regions explicitly. We then propose a novel bilateral correction network to recast the lighting of shadow regions in the illumination layer via a novel local lighting correction module, and to restore the textures conditioned on the corrected illumination layer via a novel illumination-guided texture restoration module. We further annotate pixel-wise shadow masks for the public SRD dataset, which originally contains only image pairs. Experiments on three benchmarks show that our method outperforms existing state-of-the-art shadow removal methods. Project page in: yuhaoliu7456.github.io/RRL-Net.

AAAI Conference 2024 Short Paper

THGFormer: Time-Aware Hypergraph Learning for Multimodal Social Media Popularity Prediction (Student Abstract)

  • Jienan Zhang
  • Jie Liu
  • Zhangtao Cheng
  • Xovee Xu
  • Fang Liu
  • Ting Zhong
  • Kunpeng Zhang

Social media popularity prediction of multimodal user-generated content (UGC) is a crucial task for many real-world applications. However, existing efforts are often limited by missing inter-instance correlations and UGC temporal patterns. To address these issues, we propose a novel time-aware hypergraph Transformer framework, THGFormer. It fully represents inter-instance and intra-instance relations by hypergraphs, captures the temporal dependencies with a time encoder, and enhances UGC's representations via a neighborhood knowledge aggregation. Extensive experiments conducted on two real-world datasets demonstrate that THGFormer outperforms state-of-the-art popularity prediction models across several settings.

EAAI Journal 2024 Journal Article

Uncertainty measurement for single cell RNA-seq data via Gaussian kernel: Application to unsupervised gene selection

  • Zhaowen Li
  • Jie Zhang
  • Fang Liu
  • Ching-Feng Wen

A real-valued information system (RVIS) is an information system (IS) whose information values are real numbers. If the objects, attributes and information values of a RVIS change to cells, genes and gene expression values where gene expression data is single cell RNA-seq (scRNA) data, respectively, then this RVIS is referred to as a single cell gene space ( s c g -space). Unsupervised gene selection becomes very challenging due to a lack of decision information, which is to select the optimal gene subset that can maintain learning ability without decision information. However, little research has been done on unsupervised gene selection. Uncertainty measurement is a tool of gene selection. In view of this, this paper studies uncertainty measurement in an s c g -space via Gaussian kernel and explores its application for unsupervised gene selection. In the first place, the distance between two cells in a given subspace is constructed. In the next place, the fuzzy T c o s -equivalence relation induced by this subspace is obtained employing Gaussian kernel. After that, measures of uncertainty for an s c g -space are investigated. Lastly, gene selection algorithms in an s c g -space are presented by using the proposed information entropy and information granularity. The presented algorithms are applied to clustering analyses of scRNA data. Multiple publicly available scRNA data sets are employed to evaluate the gene selection performances of the presented algorithms, while two commonly-used clustering methods, kmeans and AGNES, are utilized to obtain four metrics such as Silhouette Coefficient ( S C ), Davies–Bouldin Index ( D B I ), Fowlkes and Mallows Index ( F M I ), Normalized Mutual Information ( N M I ). The clustering results demonstrated that the presented algorithms can lower significantly the number genes selected, achieve the better S C, D B I, F M I and N M I. They also show that the presented algorithms are superior to raw data and PCA and NMF regardless of using kmeans or AGNES clustering. This also indirectly demonstrates that the granulation measure and information entropy can effectively evaluate the uncertainty of an s c g -space.

AAAI Conference 2024 Conference Paper

ViLT-CLIP: Video and Language Tuning CLIP with Multimodal Prompt Learning and Scenario-Guided Optimization

  • Hao Wang
  • Fang Liu
  • Licheng Jiao
  • Jiahao Wang
  • Zehua Hao
  • Shuo Li
  • Lingling Li
  • Puhua Chen

Pre-trained vision-language(V-L) models such as CLIP have demonstrated impressive Zero-Shot performance in many downstream tasks. Since adopting contrastive video-text pairs methods like CLIP to video tasks is limited by its high cost and scale, recent approaches focus on efficiently transferring the image-based CLIP to the video domain. A major finding is that fine-tuning the pre-trained model to achieve strong fully supervised performance leads to low zero shot, few shot, and base to novel generalization. Instead, freezing the backbone network to maintain generalization ability weakens fully supervised performance. Otherwise, no single prompt tuning branch consistently performs optimally. In this work, we proposed a multimodal prompt learning scheme that balances supervised and generalized performance. Our prompting approach contains three sections: 1) Independent prompt on both the vision and text branches to learn the language and visual contexts. 2) Inter-modal prompt mapping to ensure mutual synergy. 3) Reducing the discrepancy between the hand-crafted prompt (a video of a person doing [CLS]) and the learnable prompt, to alleviate the forgetting about essential video scenarios. Extensive validation of fully supervised, zero-shot, few-shot, base-to-novel generalization settings for video recognition indicates that the proposed approach achieves competitive performance with less commute cost.

YNICL Journal 2023 Journal Article

Force oscillations underlying precision grip in humans with lesioned corticospinal tracts

  • Charley W. Lafe
  • Fang Liu
  • Tyler W. Simpson
  • Chan Hong Moon
  • Jennifer L. Collinger
  • George F. Wittenberg
  • Michael A. Urbin

Stability of precision grip depends on the ability to regulate forces applied by the digits. Increased frequency composition and temporal irregularity of oscillations in the force signal are associated with enhanced force stability, which is thought to result from increased voluntary drive along the corticospinal tract (CST). There is limited knowledge of how these oscillations in force output are regulated in the context of dexterous hand movements like precision grip, which are often impaired by CST damage due to stroke. The extent of residual CST volume descending from primary motor cortex may help explain the ability to modulate force oscillations at higher frequencies. Here, stroke survivors with longstanding hand impairment (n = 17) and neurologically-intact controls (n = 14) performed a precision grip task requiring dynamic and isometric muscle contractions to scale and stabilize forces exerted on a sensor by the index finger and thumb. Diffusion spectrum imaging was used to quantify total white matter volume within the residual and intact CSTs of stroke survivors (n = 12) and CSTs of controls (n = 14). White matter volumes within the infarct region and an analogous portion of overlap with the CST, mirrored onto the intact side, were also quantified in stroke survivors. We found reduced ability to stabilize force and more restricted frequency ranges in force oscillations of stroke survivors relative to controls; though, more broadband, irregular output was strongly related to force-stabilizing ability in both groups. The frequency composition and temporal irregularity of force oscillations observed in stroke survivors did not correlate with maximal precision grip force, suggesting that it is not directly related to impaired force-generating capacity. The ratio of residual to intact CST volumes contained within infarct and mirrored compartments was associated with more broadband, irregular force oscillations in stroke survivors. Our findings provide insight into granular aspects of dexterity altered by corticospinal damage and supply preliminary evidence to support that the ability to modulate force oscillations at higher frequencies is explained, at least in part, by residual CST volume in stroke survivors.

AAAI Conference 2022 Conference Paper

Self-Training Multi-Sequence Learning with Transformer for Weakly Supervised Video Anomaly Detection

  • Shuo Li
  • Fang Liu
  • Licheng Jiao

Weakly supervised Video Anomaly Detection (VAD) using Multi-Instance Learning (MIL) is usually based on the fact that the anomaly score of an abnormal snippet is higher than that of a normal snippet. In the beginning of training, due to the limited accuracy of the model, it is easy to select the wrong abnormal snippet. In order to reduce the probability of selection errors, we first propose a Multi-Sequence Learning (MSL) method and a hinge-based MSL ranking loss that uses a sequence composed of multiple snippets as an optimization unit. We then design a Transformer-based MSL network to learn both video-level anomaly probability and snippet-level anomaly scores. In the inference stage, we propose to use the video-level anomaly probability to suppress the fluctuation of snippet-level anomaly scores. Finally, since VAD needs to predict the snippet-level anomaly scores, by gradually reducing the length of selected sequence, we propose a self-training strategy to gradually refine the anomaly scores. Experimental results show that our method achieves significant improvements on ShanghaiTech, UCF-Crime, and XD-Violence.

AAAI Conference 2021 Conference Paper

Dynamic Modeling Cross- and Self-Lattice Attention Network for Chinese NER

  • Shan Zhao
  • Minghao Hu
  • Zhiping Cai
  • Haiwen Chen
  • Fang Liu

Word-character lattice models have been proved to be effective for Chinese named entity recognition (NER), in which word boundary information is fused into character sequences for enhancing character representations. However, prior approaches have only used simple methods such as feature concatenation or position encoding to integrate word-character lattice information, but fail to capture fine-grained correlations in word-character spaces. In this paper, we propose DC- SAN, a Dynamic Cross- and Self-lattice Attention Network that aims to model dense interactions over word-character lattice structure for Chinese NER. By carefully combining cross-lattice and self-lattice attention modules with gated word-character semantic fusion unit, the network can explicitly capture fine-grained correlations across different spaces (e. g. , word-to-character and character-to-character), thus significantly improving model performance. Experiments on four Chinese NER datasets show that DCSAN obtains stateof-the-art results as well as efficiency compared to several competitive approaches.

IJCAI Conference 2020 Conference Paper

AttAN: Attention Adversarial Networks for 3D Point Cloud Semantic Segmentation

  • Gege Zhang
  • Qinghua Ma
  • Licheng Jiao
  • Fang Liu
  • Qigong Sun

3D point cloud semantic segmentation has attracted wide attention with its extensive applications in autonomous driving, AR/VR, and robot sensing fields. However, in existing methods, each point in the segmentation results is predicted independently from each other. This property causes the non-contiguity of label sets in three-dimensional space and produces many noisy label points, which hinders the improvement of segmentation accuracy. To address this problem, we first extend adversarial learning to this task and propose a novel framework Attention Adversarial Networks (AttAN). With high-order correlations in label sets learned from the adversarial learning, segmentation network can predict labels closer to the real ones and correct noisy results. Moreover, we design an additive attention block for the segmentation network, which is used to automatically focus on regions critical to the segmentation task by learning the correlation between multi-scale features. Adversarial learning, which explores the underlying relationship between labels in high-dimensional space, opens up a new way in 3D point cloud semantic segmentation. Experimental results on ScanNet and S3DIS datasets show that this framework effectively improves the segmentation quality and outperforms other state-of-the-art methods.

IJCAI Conference 2020 Conference Paper

Modeling Dense Cross-Modal Interactions for Joint Entity-Relation Extraction

  • Shan Zhao
  • Minghao Hu
  • Zhiping Cai
  • Fang Liu

Joint extraction of entities and their relations benefits from the close interaction between named entities and their relation information. Therefore, how to effectively model such cross-modal interactions is critical for the final performance. Previous works have used simple methods such as label-feature concatenation to perform coarse-grained semantic fusion among cross-modal instances, but fail to capture fine-grained correlations over token and label spaces, resulting in insufficient interactions. In this paper, we propose a deep Cross-Modal Attention Network (CMAN) for joint entity and relation extraction. The network is carefully constructed by stacking multiple attention units in depth to fully model dense interactions over token-label spaces, in which two basic attention units are proposed to explicitly capture fine-grained correlations across different modalities (e. g. , token-to-token and labelto-token). Experiment results on CoNLL04 dataset show that our model obtains state-of-the-art results by achieving 90. 62% F1 on entity recognition and 72. 97% F1 on relation classification. In ADE dataset, our model surpasses existing approaches by more than 1. 9% F1 on relation classification. Extensive analyses further confirm the effectiveness of our approach.

TIST Journal 2019 Journal Article

A Trust Computing-based Security Routing Scheme for Cyber Physical Systems

  • Yuxin Liu
  • Xiao Liu
  • Anfeng Liu
  • Neal N. Xiong
  • Fang Liu

Security is a pivotal issue for the development of Cyber Physical Systems (CPS). The trusted computing of CPS includes the complete protection mechanisms, such as hardware, firmware, and software, the combination of which is responsible for enforcing a system security policy. A Trust Detection-based Secured Routing (TDSR) scheme is proposed to establish security routes from source nodes to the data center under malicious environment to ensure network security. In the TDSR scheme, sensor nodes in the routing path send detection routing to identify relay nodes’ trust. And then, data packets are routed through trustworthy nodes to sink securely. In the TDSR scheme, the detection routing is executed in those nodes that have abundant energy; thus, the network lifetime cannot be affected. Performance evaluation through simulation is carried out for success of routing ratio, compromised node detection ratio, and detection routing overhead. The experiment results show that the performance can be improved in the TDSR scheme compared to previous schemes.

TIST Journal 2019 Journal Article

Edge-enabled Disaster Rescue

  • Fang Liu
  • Yeting Guo
  • Zhiping Cai
  • Nong Xiao
  • Ziming Zhao

In the aftermath of earthquakes, floods, and other disasters, photos are increasingly playing more significant roles, such as finding missing people and assessing disasters, in rescue and recovery efforts. These disaster photos are taken in real time by the crowd, unmanned aerial vehicles, and wireless sensors. However, communications equipment is often damaged in disasters, and the very limited communication bandwidth restricts the upload of photos to the cloud center, seriously impeding disaster rescue endeavors. Based on edge computing, we propose Echo, a highly time-efficient disaster rescue framework. By utilizing the computing, storage, and communication abilities of edge servers, disaster photos are preprocessed and analyzed in real time, and more specific visuals are immensely helpful for conducting emergency response and rescue. This article takes the search for missing people as a case study to show that Echo can be more advantageous in terms of disaster rescue. To greatly conserve valuable communication bandwidth, only significantly associated images are extracted and uploaded to the cloud center for subsequent facial recognition. Furthermore, an adaptive photo detector is designed to utilize the precious and unstable communication bandwidth effectively, as well as ensure the photo detection precision and recall rate. The effectiveness and efficiency of the proposed method are demonstrated by simulation experiments.

AAAI Conference 2018 Conference Paper

A Change-Detection Based Framework for Piecewise-Stationary Multi-Armed Bandit Problem

  • Fang Liu
  • Joohyun Lee
  • Ness Shroff

The multi-armed bandit problem has been extensively studied under the stationary assumption. However in reality, this assumption often does not hold because the distributions of rewards themselves may change over time. In this paper, we propose a change-detection (CD) based framework for multiarmed bandit problems under the piecewise-stationary setting, and study a class of change-detection based UCB (Upper Confidence Bound) policies, CD-UCB, that actively detects change points and restarts the UCB indices. We then develop CUSUM-UCB and PHT-UCB, that belong to the CD-UCB class and use cumulative sum (CUSUM) and Page-Hinkley Test (PHT) to detect changes. We show that CUSUM-UCB obtains the best known regret upper bound under mild assumptions. We also demonstrate the regret reduction of the CD-UCB policies over arbitrary Bernoulli rewards and Yahoo! datasets of webpage click-through rates.

YNIMG Journal 2018 Journal Article

Bayesian convolutional neural network based MRI brain extraction on nonhuman primates

  • Gengyan Zhao
  • Fang Liu
  • Jonathan A. Oler
  • Mary E. Meyerand
  • Ned H. Kalin
  • Rasmus M. Birn

Brain extraction or skull stripping of magnetic resonance images (MRI) is an essential step in neuroimaging studies, the accuracy of which can severely affect subsequent image processing procedures. Current automatic brain extraction methods demonstrate good results on human brains, but are often far from satisfactory on nonhuman primates, which are a necessary part of neuroscience research. To overcome the challenges of brain extraction in nonhuman primates, we propose a fully-automated brain extraction pipeline combining deep Bayesian convolutional neural network (CNN) and fully connected three-dimensional (3D) conditional random field (CRF). The deep Bayesian CNN, Bayesian SegNet, is used as the core segmentation engine. As a probabilistic network, it is not only able to perform accurate high-resolution pixel-wise brain segmentation, but also capable of measuring the model uncertainty by Monte Carlo sampling with dropout in the testing stage. Then, fully connected 3D CRF is used to refine the probability result from Bayesian SegNet in the whole 3D context of the brain volume. The proposed method was evaluated with a manually brain-extracted dataset comprising T1w images of 100 nonhuman primates. Our method outperforms six popular publicly available brain extraction packages and three well-established deep learning based methods with a mean Dice coefficient of 0. 985 and a mean average symmetric surface distance of 0. 220 mm. A better performance against all the compared methods was verified by statistical tests (all p-values < 10−4, two-sided, Bonferroni corrected). The maximum uncertainty of the model on nonhuman primate brain extraction has a mean value of 0. 116 across all the 100 subjects. The behavior of the uncertainty was also studied, which shows the uncertainty increases as the training set size decreases, the number of inconsistent labels in the training set increases, or the inconsistency between the training set and the testing set increases.

AAAI Conference 2018 Conference Paper

Information Directed Sampling for Stochastic Bandits With Graph Feedback

  • Fang Liu
  • Swapna Buccapatnam
  • Ness Shroff

We consider stochastic multi-armed bandit problems with graph feedback, where the decision maker is allowed to observe the neighboring actions of the chosen action. We allow the graph structure to vary with time and consider both deterministic and Erdős-Rényi random graph models. For such a graph feedback model, we first present a novel analysis of Thompson sampling that leads to tighter performance bound than existing work. Next, we propose new Information Directed Sampling based policies that are graph-aware in their decision making. Under the deterministic graph case, we establish a Bayesian regret bound for the proposed policies that scales with the clique cover number of the graph instead of the number of actions. Under the random graph case, we provide a Bayesian regret bound for the proposed policies that scales with the ratio of the number of actions over the expected number of observations per iteration. To the best of our knowledge, this is the first analytical result for stochastic bandits with random graph feedback. Finally, using numerical evaluations, we demonstrate that our proposed IDS policies outperform existing approaches, including adaptions of upper confidence bound, -greedy and Exp3 algorithms.

JMLR Journal 2018 Journal Article

Reward Maximization Under Uncertainty: Leveraging Side-Observations on Networks

  • Swapna Buccapatnam
  • Fang Liu
  • Atilla Eryilmaz
  • Ness B. Shroff

We study the stochastic multi-armed bandit (MAB) problem in the presence of side-observations across actions that occur as a result of an underlying network structure. In our model, a bipartite graph captures the relationship between actions and a common set of unknowns such that choosing an action reveals observations for the unknowns that it is connected to. This models a common scenario in online social networks where users respond to their friends' activity, thus providing side information about each other's preferences. Our contributions are as follows: 1) We derive an asymptotic lower bound (with respect to time) as a function of the bi-partite network structure on the regret of any uniformly good policy that achieves the maximum long-term average reward. 2) We propose two policies - a randomized policy; and a policy based on the well- known upper confidence bound (UCB) policies - both of which explore each action at a rate that is a function of its network position. We show, under mild assumptions, that these policies achieve the asymptotic lower bound on the regret up to a multiplicative factor, independent of the network structure. Finally, we use numerical examples on a real-world social network and a routing example network to demonstrate the benefits obtained by our policies over other existing policies. [abs] [ pdf ][ bib ] &copy JMLR 2018. ( edit, beta )

IJCAI Conference 2018 Conference Paper

UCBoost: A Boosting Approach to Tame Complexity and Optimality for Stochastic Bandits

  • Fang Liu
  • Sinong Wang
  • Swapna Buccapatnam
  • Ness Shroff

In this work, we address the open problem of finding low-complexity near-optimal multi-armed bandit algorithms for sequential decision making problems. Existing bandit algorithms are either sub-optimal and computationally simple (e. g. , UCB1) or optimal and computationally complex (e. g. , kl-UCB). We propose a boosting approach to Upper Confidence Bound based algorithms for stochastic bandits, that we call UCBoost. Specifically, we propose two types of UCBoost algorithms. We show that UCBoost(D) enjoys O(1) complexity for each arm per round as well as regret guarantee that is 1/e-close to that of the kl-UCB algorithm. We propose an approximation-based UCBoost algorithm, UCBoost(epsilon), that enjoys a regret guarantee epsilon-close to that of kl-UCB as well as O(log(1/epsilon)) complexity for each arm per round. Hence, our algorithms provide practitioners a practical way to trade optimality with computational complexity. Finally, we present numerical results which show that UCBoost(epsilon) can achieve the same regret performance as the standard kl-UCB while incurring only 1% of the computational cost of kl-UCB.

AAAI Conference 2017 Short Paper

ATSUM: Extracting Attractive Summaries for News Propagation on Microblogs

  • Fang Liu
  • Xiaojun Wan

In this paper, we investigate how to automatically extract attractive summaries for news propagation on microblogs and propose a novel system called ATSUM to achieve this goal via text attractiveness analysis. It first analyzes the sentences in a news article and automatically predict the attractiveness score of each sentence by using the support vector regression method. The predicted attractiveness scores are then incorporated into a summarization system. Experimental results on a manually labeled dataset verify the effectiveness of the proposed methods.

TIST Journal 2017 Journal Article

Energy-Efficient Mobile Video Streaming

  • Wei Zhang
  • Rui Fan
  • Yonggang Wen
  • Fang Liu

Video streaming is one of the most widely used mobile applications today, and it also accounts for a large fraction of mobile battery usage. Much of the energy consumption is for wireless data transmission and is highly correlated to network bandwidth conditions. In periods of poor connectivity, up to 90% of mobile energy can be used for wireless data transfer. In this article, we study the problem of energy-efficient mobile video streaming. We make use of the observed correlation between bandwidth and user location, and also observe that a user’s location is predictable in many situations, such as when commuting to a known destination. Based on the user’s predicted locations and bandwidth conditions, we optimize wireless transmission times to achieve high quality video playback while minimizing energy use. We propose an optimal offline algorithm for this problem, which runs in O ( Tk ) time, where T is the duration of the video and k is the size of the video buffer. We also propose LAWS, a Location AWare Streaming algorithm. LAWS learns from historical location-aware bandwidth conditions and predicts future bandwidths along a planned route to make online wireless download decisions. We evaluate LAWS using real bandwidth traces, and show that LAWS closely approximates the performance of the optimal offline algorithm, achieving 90.6% of the optimal performance on average, and 97% in certain cases. LAWS also outperforms three popular strategies used in practice by, on average, 69%, 63%, and 38%, respectively. Lastly, we show that LAWS is able to deal with noisy data and can attain the stated performance after sampling bandwidth conditions only five times.

AAAI Conference 2017 Conference Paper

Non-Additive Security Games

  • Sinong Wang
  • Fang Liu
  • Ness Shroff

Security agencies have found security games to be useful models to understand how to better protect their assets. The key practical elements in this work are: (i) the attacker can simultaneously attack multiple targets, and (ii) different targets exhibit different types of dependencies based on the assets being protected (e. g. , protection of critical infrastructure, network security, etc.). However, little is known about the computational complexity of these problems, especially when there exist dependencies among the targets. Moreover, previous security game models do not in general scale well. In this paper, we investigate a general security game where the utility function is defined on a collection of subsets of all targets, and provide a novel theoretical framework to show how to compactly represent such a game, efficiently compute the optimal (minimax) strategies, and characterize the complexity of this problem. We apply our theoretical framework to the network security game. We characterize settings under which we find a polynomial time algorithm for computing optimal strategies. In other settings we prove the problem is NP-hard and provide an approximation algorithm.

YNIMG Journal 2017 Journal Article

Oxytocin differentially alters resting state functional connectivity between amygdala subregions and emotional control networks: Inverse correlation with depressive traits

  • Monika Eckstein
  • Sebastian Markett
  • Keith M. Kendrick
  • Beate Ditzen
  • Fang Liu
  • Rene Hurlemann
  • Benjamin Becker

The hypothalamic neuropeptide oxytocin (OT) has received increasing attention for its role in modulating social-emotional processes across species. Previous studies on using intranasal-OT in humans point to a crucial engagement of the amygdala in the observed neuromodulatory effects of OT under task and rest conditions. However, the amygdala is not a single homogenous structure, but rather a set of structurally and functionally heterogeneous nuclei that show distinct patterns of connectivity with limbic and frontal emotion-processing regions. To determine potential differential effects of OT on functional connectivity of the amygdala subregions, 79 male participants underwent resting-state fMRI following randomized intranasal-OT or placebo administration. In line with previous studies OT increased the connectivity of the total amygdala with dorso-medial prefrontal regions engaged in emotion regulation. In addition, OT enhanced coupling of the total amygdala with cerebellar regions. Importantly, OT differentially altered the connectivity of amygdala subregions with distinct up-stream cortical nodes, particularly prefrontal/parietal, and cerebellar down-stream regions. OT-induced increased connectivity with cerebellar regions were largely driven by effects on the centromedial and basolateral subregions, whereas increased connectivity with prefrontal regions were largely mediated by right superficial and basolateral subregions. OT decreased connectivity of the centromedial subregions with core hubs of the emotional face processing network in temporal, occipital and parietal regions. Preliminary findings suggest that effects on the superficial amygdala-prefrontal pathway were inversely associated with levels of subclinical depression, possibly indicating that OT modulation may be blunted in the context of increased pathological load. Together, the present findings suggest a subregional-specific modulatory role of OT on amygdala-centered emotion processing networks in humans.

YNICL Journal 2016 Journal Article

Changes of grey matter volume in first-episode drug-naive adult major depressive disorder patients with different age-onset

  • Zonglin Shen
  • Yuqi Cheng
  • Shuran Yang
  • Nan Dai
  • Jing Ye
  • Xiaoyan Liu
  • Jin Lu
  • Na Li

OBJECTIVE: Little is known about the pathological mechanism of early adult onset depression (EOD) and later adult onset depression (LOD). We seek to determine whether grey matter volume (GMV) change in EOD and LOD are different, which could also delineate EOD and LOD. METHODS: In present study, 147 first-episode, drug-naive patients with major depressive disorder (MDD), age between 18 and 45, were divided into two groups on the basis of age of MDD onset: the early adult onset group (age 18-29) and the later adult onset group (age 30-44), and a total of 130 gender-, and age-, matched healthy controls (HC) were also divided into two groups which fit for each patient group. Magnetic resonance imaging was conducted on all subjects. The voxel-based morphometry (VBM) approach was employed to analyze the images. RESULTS: Widespread abnormalities of GMV throughout parietal, temporal, limbic regions, occipital cortex and cerebellum were observed in MDD patients. Compare to young HC, reduced GMV in right fusiform gyrus, right middle temporal gyrus, vermis III and increased GMV in right middle occipital gyrus were seen in the EOD group. In contrast, relative to old HC, decreased GMV in the right hippocampus and increased GMV in the left middle temporal gyrus were observed in the LOD group. Compared to the LOD group, the EOD group had smaller GMV in right posterior cingulate cortex. There was no significant correlation between GMV of the right posterior cingulate cortex and the score of the depression rating scale in patients group. CONCLUSIONS: The GMV of the brain areas that were related to mood regulation was decreased in the first-episode, drug-naive adult patients with MDD. Adult patients with EOD and LOD exhibited different GMV changes relative to each age-matched comparison group, suggesting depressed adult patients with different age-onset might have different pathological mechanism.

IJCAI Conference 2013 Conference Paper

Improving Question Retrieval in Community Question Answering Using World Knowledge

  • Guangyou Zhou
  • Yang Liu
  • Fang Liu
  • Daojian Zeng
  • Jun Zhao

Community question answering (cQA), which provides a platform for people with diverse background to share information and knowledge, has become an increasingly popular research topic. In this paper, we focus on the task of question retrieval. The key problem of question retrieval is to measure the similarity between the queried questions and the historical questions which have been solved by other users. The traditional methods measure the similarity based on the bag-of-words (BOWs) representation. This representation neither captures dependencies between related words, nor handles synonyms or polysemous words. In this work, we first propose a way to build a concept thesaurus based on the semantic relations extracted from the world knowledge of Wikipedia. Then, we develop a unified framework to leverage these semantic relations in order to enhance the question similarity in the concept space. Experiments conducted on a real cQA data set show that with the help of Wikipedia thesaurus, the performance of question retrieval is improved as compared to the traditional methods.

EAAI Journal 2012 Journal Article

Immune optimization algorithm for solving joint call admission control problem in next-generation wireless network

  • Si-Feng Zhu
  • Fang Liu
  • Yu-tao Qi
  • Zheng-yi Chai
  • Jian-she Wu

The integration of radio access networks with different radio access technologies (RATs) is one of the remarkable characteristics of the next-generation wireless networks (NGWNs). In NGWN, the users with multi-network interface terminals should be able to select independently radio access network to obtain the best service. Therefore, joint call admission control (JCAC) schemes are required to select the most appropriate radio access network (RAN) for incoming calls. We propose an immune algorithm-based JCAC (IA-JCAC) scheme with users centric in order to enhance user's satisfaction. However, JCAC algorithms with users centric can lead to highly unbalanced traffic load among the available RANs in NGWN because users act independently, and most of them may prefer to be connected through a particular RAN. Highly unbalanced traffic load in NGWN will result in high overall call blocking/dropping probability and poor radio result utilization. To solve this problem, we employ dynamic pricing for balancing traffic load among available RANs in heterogeneous wireless networks where users' preferences are considered in decision-making on RAT selection. The proposed IA-based JCAC scheme is compared with another scheme that does not use the dynamic pricing on the performance. The simulation result shows the effectiveness of the proposed IA-JCAC scheme is improved significantly.

v2026.09.13