Arrow Research search

Author name cluster

Fang Deng

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAAI Conference 2026 Conference Paper

Hierarchical Reinforcement Learning with Topology-Aware Exploration Framework for Multi-path Commodity Flow Problem

  • Jingchen Jiang
  • Xuan Zhou
  • Jiayuan Li
  • Geng Han
  • Xiang Shi
  • Fang Deng

The multi-path commodity flow problem (MPCFP) is crucial for ensuring reliable and high-speed data transmission in communication networks. However, existing studies that employ pre-generated routing paths neglect real-time load state and the coupling among decisions, thus hindering the achievement of high-quality solutions. To overcome this, we propose Hierarchical Reinforcement Learning with Topology-Aware Exploration (HRL-TAE), which is the first fully end-to-end framework that dynamically produces high-quality solutions based on real-time network states. HRL-TAE integrates an exploration mechanism and utilizes the State Transition Guiding List (STGL) to guide state transitions, thereby transforming topology exploration into a Markov decision process. Guided by STGL, two closely coupled layers in HRL-TAE, that is, the path construct layer and the ratio allocate layer, construct multiple subpaths for each flow and allocate traffic ratios among them. Subsequently, adaptive constraint-driven masks exclude infeasible actions during decision making, thereby guaranteeing that all constraints are satisfied. We also adopt a tailored training approach to obtain accurate gradient estimates and improve training efficiency. Simulations and real-world experiments demonstrate that HRL-TAE achieves superior performance.

AAAI Conference 2025 Conference Paper

From Coarse to Fine: A Matching and Alignment Framework for Unsupervised Cross-View Geo-Localization

  • Xueyi Wang
  • Lele Zhang
  • Zheng Fan
  • Yang Liu
  • Chen Chen
  • Fang Deng

Cross-view geo-localization aims at determining the geographic location of a query image by matching the reference images. The matching pairs can be captured from diverse perspectives, such as those from satellites and drones. Most existing methods are supervised that require input of location-labeled images or matched and unmatched image pairs for training, resulting in high labor costs. Moreover, current unsupervised methods perform instances matching directly between different perspectives with dramatic discrepancies, resulting in poor performance. To address these issues, this paper proposes a novel matching and alignment framework from coarse instance-cluster level to fine intermediate instance level for unsupervised cross-view geo-localization. We first introduces cluster-based contrastive learning, assigning pseudo-labels to the instances and generate clusters within each view. Then we design a cross-view location alignment module that fully exploits the feature relationships between instances and clusters for intra- and inter-views. Finally, we design an intermediate state transition module that facilitates further alignment between views by constructing intermediate states and bringing both views closer to the intermediate domain simultaneously. Extensive experiments demonstrate that our method surpasses state-of-the-art unsupervised cross-view geo-localization methods and even achieves comparable performance to state-of-the-art supervised methods.

ICML Conference 2025 Conference Paper

In-Context Adaptation to Concept Drift for Learned Database Operations

  • Jiaqi Zhu 0002
  • Shaofeng Cai
  • Yanyan Shen
  • Gang Chen 0001
  • Fang Deng
  • Beng Chin Ooi

Machine learning has demonstrated transformative potential for database operations, such as query optimization and in-database data analytics. However, dynamic database environments, characterized by frequent updates and evolving data distributions, introduce concept drift, which leads to performance degradation for learned models and limits their practical applicability. Addressing this challenge requires efficient frameworks capable of adapting to shifting concepts while minimizing the overhead of retraining or fine-tuning. In this paper, we propose FLAIR, an online adaptation framework that introduces a new paradigm called in-context adaptation for learned database operations. FLAIR leverages the inherent property of data systems, i. e. , immediate availability of execution results for predictions, to enable dynamic context construction. By formalizing adaptation as $f: (\mathbf{x} | \mathcal{C}_t) \to \mathbf{y}$, with $\mathcal{C}_t$ representing a dynamic context memory, FLAIR delivers predictions aligned with the current concept, eliminating the need for runtime parameter optimization. To achieve this, FLAIR integrates two key modules: a Task Featurization Module for encoding task-specific features into standardized representations, and a Dynamic Decision Engine, pre-trained via Bayesian meta-training, to adapt seamlessly using contextual information at runtime. Extensive experiments across key database tasks demonstrate that FLAIR outperforms state-of-the-art baselines, achieving up to $5. 2\times$ faster adaptation and reducing error by 22. 5% for cardinality estimation.

NeurIPS Conference 2025 Conference Paper

Learning CAD Modeling Sequences via Projection and Part Awareness

  • Yang Liu
  • Daxuan Ren
  • Yijie Ding
  • Jianmin Zheng
  • Fang Deng

This paper presents PartCAD, a novel framework for reconstructing CAD modeling sequences directly from point clouds by projection-guided, part-aware geometry reasoning. It consists of (1) an autoregressive approach that decomposes point clouds into part-aware latent representations, serving as interpretable anchors for CAD generation; (2) a projection guidance module that provides explicit cues about underlying design intent via triplane projections; and (3) a non-autoregressive decoder to generate sketch-extrusion parameters in a single forward pass, enabling efficient and structurally coherent CAD instruction synthesis. By bridging geometric signals and semantic understanding, PartCAD tackles the challenge of reconstructing editable CAD models—capturing underlying design processes—from 3D point clouds. Extensive experiments show that PartCAD significantly outperforms existing methods for CAD instruction generation in both accuracy and robustness. The work sheds light on part-driven reconstruction of interpretable CAD models, opening new avenues in reverse engineering and CAD automation.

YNIMG Journal 2025 Journal Article

Online and in-person collaborative writing have similar benefits but different costs

  • Hengyue Ran
  • Qi Li
  • Yuanyuan Li
  • Fang Deng
  • Yafeng Pan

With the rapid rise of online education, collaborative learning is no longer confined to physical classrooms. Yet, it remains unclear whether online collaboration, especially with or without visual cues, can support the same cognitive and neural processes as in-person collaboration. This study used multimodal learning analytics to compare collaboration processes and inter-brain synchronization (IBS) under three conditions: in-person, online with camera on, and online with camera off. Seventy-seven learner dyads completed a 28-minute collaborative writing task while their brain activity was recorded simultaneously using functional near-infrared spectroscopy (fNIRS). Across all three conditions, collaborative learning significantly improved outcomes. In-person and online (camera on) learners showed comparable IBS in the middle temporal gyrus. However, camera-on learners displayed more frequent higher-order behaviors (e.g., monitoring, questioning, mutual understanding, argument building) and greater dorsolateral prefrontal cortex activation, reflecting increased executive control demands. In contrast, camera-off learners achieved learning gains but engaged in less information exchange, emphasized mutual understanding and collaborative planning, and exhibited markedly lower IBS. Together, these findings indicate that while both in-person and online collaboration can yield similar levels of achievement, their cognitive costs differ: in-person collaboration is more efficient, whereas online collaboration requires additional regulation and cognitive effort. The absence of visual cues further constrains information sharing and social interaction, undermining IBS. These insights help explain the mechanisms that shape collaborative learning across contexts and offer guidance for designing more effective online learning environments.

YNIMG Journal 2024 Journal Article

Clinical characteristics of post-stroke basal ganglia aphasia and the study of language-related white matter tracts based on diffusion spectrum imaging

  • Yue Han
  • Yuanyuan Jing
  • Xuewei Li
  • Hongwei Zhou
  • Fang Deng

BACKGROUND: Stroke often damages the basal ganglia, leading to atypical and transient aphasia, indicating that post-stroke basal ganglia aphasia (PSBGA) may be related to different anatomical structural damage and functional remodeling rehabilitation mechanisms. The basal ganglia contain dense white matter tracts (WMTs). Hence, damage to the functional tract may be an essential anatomical structural basis for the development of PSBGA. METHODS: We first analyzed the clinical characteristics of PSBGA in 28 patients and 15 healthy controls (HCs) using the Western Aphasia Battery and neuropsychological test batteries. Moreover, we investigated white matter injury during the acute stage using diffusion magnetic resonance imaging scans for differential tractography. Finally, we used multiple regression models in correlation tractography to analyze the relationship between various language functions and quantitative anisotropy (QA) of WMTs. RESULTS: Compared with HCs, patients with PSBGA showed lower scores for fluency, comprehension (auditory word recognition and sequential commands), naming (object naming and word fluency), reading comprehension of sentences, Mini-Mental State Examination, and Montreal Cognitive Assessment, along with increased scores in Hamilton Anxiety Scale-17 and Hamilton Depression Scale-17 within 7 days after stroke onset (P < 0.05). Differential tractography revealed that patients with PSBGA had damaged fibers, including in the body fibers of the corpus callosum, left cingulum bundles, left parietal aslant tracts, bilateral superior longitudinal fasciculus II, bilateral thalamic radiation tracts, left fornix, corpus callosum tapetum, and forceps major, compared with HCs (FDR < 0.02). Correlation tractography highlighted that better comprehension was correlated with a higher QA of the left inferior fronto-occipital fasciculus (IFOF), corpus callosum forceps minor, and left extreme capsule (FDR < 0.0083). Naming was positively associated with the QA of the left IFOF, forceps minor, left arcuate fasciculus, and uncinate fasciculus (UF) (FDR < 0.0083). Word fluency of naming was also positively associated with the QA of the forceps minor, left IFOF, and thalamic radiation tracts (FDR < 0.0083). Furthermore, reading was positively correlated with the QA of the forceps minor, left IFOF, and UF (FDR < 0.0083). CONCLUSION: PSBGA is primarily characterized by significantly impaired word fluency of naming and preserved repetition abilities, as well as emotional and cognitive dysfunction. Damaged limbic pathways, dorsally located tracts in the left hemisphere, and left basal ganglia pathways are involved in PSBGA pathogenesis. The results of connectometry analysis further refine the current functional localization model of higher-order neural networks associated with language functions.

IROS Conference 2024 Conference Paper

STL-SLAM: A Structured-Constrained RGB-D SLAM Approach to Texture-Limited Environments

  • Juan Dong
  • Maobin Lu
  • Chen Chen 0044
  • Fang Deng
  • Jie Chen 0003

Most RGB-D-based SLAM methods assume texture-rich environments, making them susceptible to significant tracking errors or complete failures in the absence of texture features. Moreover, many existing methods encounter substantial rotation estimation errors, leading to long-term drift in tracking. This paper proposes a novel structured-constrained RGB-D SLAM method (STL-SLAM) for texture-limited environments. Compared to the existing methods, STL-SLAM can deal with environments without abundant texture information and significantly reduce long-term drift caused by rotation estimation errors. We assess the distribution complexity of pixels in an image by calculating the information entropy and pre-processing accordingly. We also present an efficient Manhattan Frames (MF) detection strategy based on orthogonal planes and lines. If MF is detected, we decouple rotation and translation, estimate drift-free rotation based on the Manhattan World (MW) coordinate system, and then estimate translation by minimizing the re-projection error of point, line, and plane features. In non-Manhattan Frames, the 6-DoF pose estimation is performed holistically, with the incorporation of structural constraints of parallel and perpendicular planes, as well as parallel and vertical lines, into the optimization process. Finally, we evaluate our method on public datasets and in real-world environments, which shows that our proposed method achieves superior performance compared to its counterparts.

EAAI Journal 2023 Journal Article

LOSN: Lightweight ore sorting networks for edge device environment

  • Yang Liu
  • Xueyi Wang
  • Zelin Zhang
  • Fang Deng

Vision-based intelligent ore sorting technology has been widely applied in current mining production, a trend further facilitated by the emergence of deep learning. However, most available implementations are still based on image classification, i. e. , dividing the overall sorting task into two processes: classification and localization, without end-to-end integration. Meanwhile, harsh sorting scenarios make edge computing devices the primary candidate for model deployment, with more stringent limitations for model size, computational complexity, and inference speed. Therefore, this study proposes to integrate the operating processes to locate and classify the ores particles simultaneously. The lightweight structures, attention mechanisms, and multi-scale feature fusion strategies are applied in the architecture design to meet the deployment requirements of edge device environments and achieve a preferred accuracy–efficiency tradeoff, which leads to a new lightweight ore sorting networks called LOSN. In the case study, LOSN has the highest accuracy in multi-type and multi-class ore sorting tasks (78. 87% and 80. 64% in the gas coal and anthracite dataset, respectively) with fewer parameters (5. 970M), lower GFLOPs (6. 829G) and higher FPS (89. 92), which is superior to commonly used high-performance object detection architectures (e. g. , Yolo series, EfficientDet, Faster-RCNN, and CenterNet). Grad-CAM visualizations also demonstrate the feature extraction capability of LOSN.

NeurIPS Conference 2023 Conference Paper

Triangulation Residual Loss for Data-efficient 3D Pose Estimation

  • Jiachen Zhao
  • Tao Yu
  • Liang An
  • Yipeng Huang
  • Fang Deng
  • Qionghai Dai

This paper presents Triangulation Residual loss (TR loss) for multiview 3D pose estimation in a data-efficient manner. Existing 3D supervised models usually require large-scale 3D annotated datasets, but the amount of existing data is still insufficient to train supervised models to achieve ideal performance, especially for animal pose estimation. To employ unlabeled multiview data for training, previous epipolar-based consistency provides a self-supervised loss that considers only the local consistency in pairwise views, resulting in limited performance and heavy calculations. In contrast, TR loss enables self-supervision with global multiview geometric consistency. Starting from initial 2D keypoint estimates, the TR loss can fine-tune the corresponding 2D detector without 3D supervision by simply minimizing the smallest singular value of the triangulation matrix in an end-to-end fashion. Our method achieves the state-of-the-art 25. 8mm MPJPE and competitive 28. 7mm MPJPE with only 5\% 2D labeled training data on the Human3. 6M dataset. Experiments on animals such as mice demonstrate our TR loss's data-efficient training ability.

v2026.09.13