Arrow Research search

Author name cluster

Mohammed Bennamoun

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

18 papers
2 author rows

Possible papers

18

JBHI Journal 2026 Journal Article

RLAD: A Reliable Hippo-Guided Multi-Task Model for Alzheimer's Disease Diagnosis

  • Zhenxin Lei
  • Wenjing Zhu
  • Jiale Liu
  • Cong Hua
  • Johann Li
  • Syed Afaq Ali Shah
  • Liang Zhang
  • Mohammed Bennamoun

Early diagnosis of Alzheimer's disease (AD) is crucial for its prevention, and hippocampal atrophy is a significant lesion for early diagnosis. The current DL-based AD diagnosis methods only focus on either AD classification or hippocampus segmentation independently, neglecting the correlation between the two tasks and lacking pathological interpretability. To address this issue, we propose a Reliable Hippo-guided Learning model for Alzheimer's Disease diagnosis (RLAD), which employs multi-task learning for AD classification as a main task supplemented by hippocampus segmentation. More specifically, our model consists of 1) a hybrid shared features encoder that encodes local and global information in MRI to enhance the model's ability to learn discriminative features; 2) Task Specific Decoders to accomplish AD classification and hippocampus segmentation; and 3) Task Coordination module to correlate the two tasks and guide the classification task to focus on the hippocampus area. Our proposed RLAD model is evaluated on MRI scans of 1631 subjects from three independent datasets, including ADNI-1, ADNI-2, and HarP. Our extensive experimental results demonstrate that the proposed model significantly improves the performance of AD classification and hippocampus segmentation with strong generalization capabilities.

NeurIPS Conference 2025 Conference Paper

Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time Adaptation

  • Jingmin Zhu
  • Anqi Zhu
  • Hossein Rahmani
  • Jun Liu
  • Mohammed Bennamoun
  • Qiuhong Ke

We introduce Skeleton-Cache, the first training-free test-time adaptation framework for skeleton-based zero-shot action recognition (SZAR), aimed at improving model generalization to unseen actions during inference. Skeleton-Cache reformulates inference as a lightweight retrieval process over a non-parametric cache that stores structured skeleton representations, combining both global and fine-grained local descriptors. To guide the fusion of descriptor-wise predictions, we leverage the semantic reasoning capabilities of large language models (LLMs) to assign class-specific importance weights. By integrating these structured descriptors with LLM-guided semantic priors, Skeleton-Cache dynamically adapts to unseen actions without any additional training or access to training data. Extensive experiments on NTU RGB+D 60/120 and PKU-MMD II demonstrate that Skeleton-Cache consistently boosts the performance of various SZAR backbones under both zero-shot and generalized zero-shot settings. The code is publicly available at https: //github. com/Alchemist0754/Skeleton-Cache.

ICML Conference 2025 Conference Paper

GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

  • Akashah Shabbir
  • Mohammed Zumri
  • Mohammed Bennamoun
  • Fahad Shahbaz Khan
  • Salman H. Khan 0001

Recent advances in large multimodal models (LMMs) have recognized fine-grained grounding as an imperative factor of visual understanding and dialogue. However, the benefits of such representation in LMMs are limited to the natural image domain, and these models perform poorly for remote sensing (RS). The distinct overhead viewpoint, scale variation, and presence of small objects in high-resolution RS imagery present a unique challenge in region-level comprehension. Moreover, the development of the grounding conversation capability of LMMs within RS is hindered by the lack of granular, RS domain-specific grounded data. Addressing these limitations, we propose GeoPixel - the first end-to-end high-resolution RS-LMM that supports pixel-level grounding. This capability allows fine-grained visual perception by generating interleaved masks in conversation. GeoPixel supports up to 4K HD resolution in any aspect ratio, ideal for high-precision RS image analysis. To support the grounded conversation generation (GCG) in RS imagery, we curate a visually grounded dataset GeoPixelD through a semi-automated pipeline that utilizes set-of-marks prompting and spatial priors tailored for RS data to methodically control the data generation process. GeoPixel demonstrates superior performance in pixel-level comprehension, surpassing existing LMMs in both single-target and multi-target segmentation tasks. Our methodological ablation studies validate the effectiveness of each component in the overall architecture. Our code and data will be publicly released.

EAAI Journal 2025 Journal Article

Memory guided representation learning for cross-domain face anti-spoofing

  • Pengchao Deng
  • Yanhui Zhou
  • Zhiheng Fu
  • Jian Liu
  • Shengjun Xu
  • Chenyang Ge
  • Farid Boussaid
  • Mohammed Bennamoun

Addressing generalized Face Anti-Spoofing is challenging due to the wide variety of spoofing techniques, variations in environmental conditions, and the diversity of devices used to capture images. Most approaches to enhancing the generalization ability of systems often manipulate image statistics to normalize images into a uniform representation space. However, this approach can restrict the capacity of the system to represent images accurately because image statistics vary significantly across different domains, each with unique characteristics. Recognizing that each source domain possesses distinct characteristics, we introduce an innovative approach based on memory guided representation learningto represent these characteristics in separate latent spaces. Specifically, we present a dual component framework comprising a Memory Guided Clustering Representation (MGCR)generator and a Memory Guided Mapping Representation (MGMR)classifier. Additionally, we create a memory bank filled with template and meta features, which are refined over time using a momentum update mechanism. The MGCR component employs clustering to allow new, unseen deep features close to the most similar template feature, thereby creating generalized clues for the generator. Meanwhile, the MGMR process leverages meta features to dynamically represent new domain spaces through linear combinations, bridging the divide between known and unknown domains. Our experimental results show that the proposed method significantly enhances the effectiveness under cross-domain scenarios, outperforming existing techniques. Especially in the “Leave One Out”setting, its average Half Total Error Rate exceeds that of other methods, reaching less than 10%.

NeurIPS Conference 2025 Conference Paper

Watch and Listen: Understanding Audio-Visual-Speech Moments with Multimodal LLM

  • Zinuo Li
  • Xian Zhang
  • Yongxin Guo
  • Mohammed Bennamoun
  • Farid Boussaid
  • Girish Dwivedi
  • Luqi Gong
  • Qiuhong Ke

Humans naturally understand moments in a video by integrating visual and auditory cues. For example, localizing a scene in the video like “A scientist passionately speaks on wildlife conservation as dramatic orchestral music plays, with the audience nodding and applauding” requires simultaneous processing of visual, audio, and speech signals. However, existing models often struggle to effectively fuse and interpret audio information, limiting their capacity for comprehensive video temporal understanding. To address this, we present TriSense, a triple-modality large language model designed for holistic video temporal understanding through the integration of visual, audio, and speech modalities. Central to TriSense is a Query-Based Connector that adaptively reweights modality contributions based on the input query, enabling robust performance under modality dropout and allowing flexible combinations of available inputs. To support TriSense's multimodal capabilities, we introduce TriSense-2M, a high-quality dataset of over 2 million curated samples generated via an automated pipeline powered by fine-tuned LLMs. TriSense-2M includes long-form videos and diverse modality combinations, facilitating broad generalization. Extensive experiments across multiple benchmarks demonstrate the effectiveness of TriSense and its potential to advance multimodal video analysis.

EAAI Journal 2024 Journal Article

Enhancing security in X-ray baggage scans: A contour-driven learning approach for abnormality classification and instance segmentation

  • Abdelfatah Ahmed
  • Divya Velayudhan
  • Taimur Hassan
  • Mohammed Bennamoun
  • Ernesto Damiani
  • Naoufel Werghi

The task of automatically identifying hazardous items in luggage through X-ray scans is highly crucial, yet immensely complex. This method presents an improvement over the traditional, labor-intensive, time-consuming, and error-prone manual evaluation of X-ray images. In the present day, an enormous volume of luggage is continually passing through airports, seaports, and land ports, making these advancements exceedingly vital. Nevertheless, numerous existing solutions grapple with the issue of data imbalance, as the frequency of dangerous items is relatively low, adversely affecting the efficiency of networks. In this study, we introduce a solution that merges a contour-driven learned model with a unique loss function known as Balanced Affinity Loss. This function equitably distributes the focus of training models towards less represented items. We tested the proposed system using three public luggage X-ray datasets, where it exceeded the performance of the most advanced methods by 34. 1%, 8. 64%, and 10. 04% in terms of intersection-over-union for instance segmentation. Likewise, the introduced system registered improvements of 9. 11% and 3. 66% for abnormality classification.

NeurIPS Conference 2024 Conference Paper

Referring Human Pose and Mask Estimation In the Wild

  • Bo Miao
  • Mingtao Feng
  • Zijie Wu
  • Mohammed Bennamoun
  • Yongsheng Gao
  • Ajmal Mian

We introduce Referring Human Pose and Mask Estimation (R-HPM) in the wild, where either a text or positional prompt specifies the person of interest in an image. This new task holds significant potential for human-centric applications such as assistive robotics and sports analysis. In contrast to previous works, R-HPM (i) ensures high-quality, identity-aware results corresponding to the referred person, and (ii) simultaneously predicts human pose and mask for a comprehensive representation. To achieve this, we introduce a large-scale dataset named RefHuman, which substantially extends the MS COCO dataset with additional text and positional prompt annotations. RefHuman includes over 50, 000 annotated instances in the wild, each equipped with keypoint, mask, and prompt annotations. To enable prompt-conditioned estimation, we propose the first end-to-end promptable approach named UniPHD for R-HPM. UniPHD extracts multimodal representations and employs a proposed pose-centric hierarchical decoder to process (text or positional) instance queries and keypoint queries, producing results specific to the referred person. Extensive experiments demonstrate that UniPHD produces quality results based on user-friendly prompts and achieves top-tier performance on RefHuman val and MS COCO val2017.

ICRA Conference 2023 Conference Paper

Reinforced Learning for Label-Efficient 3D Face Reconstruction

  • Hoda Mohaghegh
  • Hossein Rahmani 0001
  • Hamid Laga
  • Farid Boussaïd
  • Mohammed Bennamoun

3D face reconstruction plays a major role in many human-robot interaction systems, from automatic face authentication to human-computer interface-based entertainment. To improve robustness against occlusions and noise, 3D face reconstruction networks are often trained on a set of in-the-wild face images preferably captured along different viewpoints of the subject. However, collecting the required large amounts of 3D annotated face data is expensive and time-consuming. To address the high annotation cost and due to the importance of training on a useful set, we propose an Active Learning (AL) framework that actively selects the most informative and representative samples to be labeled. To the best of our knowledge, this paper is the first work on tackling active learning for 3D face reconstruction to enable a label-efficient training strategy. In particular, we propose a Reinforcement Active Learning approach in conjunction with a clustering-based pooling strategy to select informative view-points of the subjects. Experimental results on 300W-LP and AFLW2000 datasets demonstrate that our proposed method is able to 1) efficiently select the most influencing view-points for labeling and outperforms several baseline AL techniques and 2) further improve the performance of a 3D Face Reconstruction network trained on the full dataset.

NeurIPS Conference 2023 Conference Paper

UE4-NeRF:Neural Radiance Field for Real-Time Rendering of Large-Scale Scene

  • Jiaming Gu
  • Minchao Jiang
  • Hongsheng Li
  • Xiaoyuan Lu
  • Guangming Zhu
  • Syed Afaq Ali Shah
  • Liang Zhang
  • Mohammed Bennamoun

Neural Radiance Fields (NeRF) is a novel implicit 3D reconstruction method that shows immense potential and has been gaining increasing attention. It enables the reconstruction of 3D scenes solely from a set of photographs. However, its real-time rendering capability, especially for interactive real-time rendering of large-scale scenes, still has significant limitations. To address these challenges, in this paper, we propose a novel neural rendering system called UE4-NeRF, specifically designed for real-time rendering of large-scale scenes. We partitioned each large scene into different sub-NeRFs. In order to represent the partitioned independent scene, we initialize polygonal meshes by constructing multiple regular octahedra within the scene and the vertices of the polygonal faces are continuously optimized during the training process. Drawing inspiration from Level of Detail (LOD) techniques, we trained meshes of varying levels of detail for different observation levels. Our approach combines with the rasterization pipeline in Unreal Engine 4 (UE4), achieving real-time rendering of large-scale scenes at 4K resolution with a frame rate of up to 43 FPS. Rendering within UE4 also facilitates scene editing in subsequent stages. Furthermore, through experiments, we have demonstrated that our method achieves rendering quality comparable to state-of-the-art approaches. Project page: https: //jamchaos. github. io/UE4-NeRF/.

NeurIPS Conference 2022 Conference Paper

Active-Passive SimStereo - Benchmarking the Cross-Generalization Capabilities of Deep Learning-based Stereo Methods

  • Laurent Jospin
  • Allen Antony
  • Lian Xu
  • Hamid Laga
  • Farid Boussaid
  • Mohammed Bennamoun

In stereo vision, self-similar or bland regions can make it difficult to match patches between two images. Active stereo-based methods mitigate this problem by projecting a pseudo-random pattern on the scene so that each patch of an image pair can be identified without ambiguity. However, the projected pattern significantly alters the appearance of the image. If this pattern acts as a form of adversarial noise, it could negatively impact the performance of deep learning-based methods, which are now the de-facto standard for dense stereo vision. In this paper, we propose the Active-Passive SimStereo dataset and a corresponding benchmark to evaluate the performance gap between passive and active stereo images for stereo matching algorithms. Using the proposed benchmark and an additional ablation study, we show that the feature extraction and matching modules of a selection of twenty selected deep learning-based stereo matching methods generalize to active stereo without a problem. However, the disparity refinement modules of three of the twenty architectures (ACVNet, CascadeStereo, and StereoNet) are negatively affected by the active stereo patterns due to their reliance on the appearance of the input images.

NeurIPS Conference 2018 Conference Paper

Attention in Convolutional LSTM for Gesture Recognition

  • Liang Zhang
  • Guangming Zhu
  • Lin Mei
  • Peiyi Shen
  • Syed Afaq Ali Shah
  • Mohammed Bennamoun

Convolutional long short-term memory (LSTM) networks have been widely used for action/gesture recognition, and different attention mechanisms have also been embedded into the LSTM or the convolutional LSTM (ConvLSTM) networks. Based on the previous gesture recognition architectures which combine the three-dimensional convolution neural network (3DCNN) and ConvLSTM, this paper explores the effects of attention mechanism in ConvLSTM. Several variants of ConvLSTM are evaluated: (a) Removing the convolutional structures of the three gates in ConvLSTM, (b) Applying the attention mechanism on the input of ConvLSTM, (c) Reconstructing the input and (d) output gates respectively with the modified channel-wise attention mechanism. The evaluation results demonstrate that the spatial convolutions in the three gates scarcely contribute to the spatiotemporal feature fusion, and the attention mechanisms embedded into the input and output gates cannot improve the feature fusion. In other words, ConvLSTM mainly contributes to the temporal fusion along with the recurrent steps to learn the long-term spatiotemporal features, when taking as input the spatial or spatiotemporal features. On this basis, a new variant of LSTM is derived, in which the convolutional structures are only embedded into the input-to-state transition of LSTM. The code of the LSTM variants is publicly available.

IJCAI Conference 2017 Conference Paper

Learning deep structured network for weakly supervised change detection

  • Salman Khan
  • Xuming He
  • Fatih Porikli
  • Mohammed Bennamoun
  • Ferdous Sohel
  • Roberto Togneri

Conventional change detection methods require a large number of images to learn background models or depend on tedious pixel-level labeling by humans. In this paper, we present a weakly supervised approach that needs only image-level labels to simultaneously detect and localize changes in a pair of images. To this end, we employ a deep neural network with DAG topology to learn patterns of change from image-level labeled training data. On top of the initial CNN activations, we define a CRF model to incorporate the local differences and context with the dense connections between individual pixels. We apply a constrained mean-field algorithm to estimate the pixel-level labels, and use the estimated labels to update the parameters of the CNN in an iterative EM framework. This enables imposing global constraints on the observed foreground probability mass function. Our evaluations on four benchmark datasets demonstrate superior detection and localization performance.

ICRA Conference 2016 Conference Paper

Simultaneous dense scene reconstruction and object labeling

  • Umar Asif
  • Mohammed Bennamoun
  • Ferdous Ahmed Sohel

This paper presents an efficient system for simultaneous dense scene reconstruction and object labeling in real-world environments (captured with an RGB-D sensor). The proposed system starts with the generation of object proposals in the scene. It then tracks spatio-temporally consistent object proposals across multiple frames and produces a dense reconstruction of the scene. In parallel, the proposed system uses an efficient inference algorithm, where object class probabilities are computed at an object-level and fused into a voxel-based prediction hypothesis modeled on the voxels of the reconstructed scene. Our extensive experiments using challenging RGB-D object and scene datasets, and live video streams from Microsoft Kinect show that the proposed system achieved competitive 3D scene reconstruction and object labeling results compared to the state-of-the-art methods.

IROS Conference 2015 Conference Paper

Discriminative feature learning for efficient RGB-D object recognition

  • Umar Asif
  • Mohammed Bennamoun
  • Ferdous Ahmed Sohel

This paper presents an efficient approach to recognize objects captured with an RGB-D sensor. The proposed approach uses a Bag-of-Words (BOW) model to learn feature representations from raw RGB-D point clouds in a weakly supervised manner. To this end, we introduce a novel method based on randomized clustering trees to learn visual vocabularies which are fast to compute and more discriminative compared to the vocabularies generated by classical methods such as k-means. We show that, when combined with standard spatial pooling strategies, our proposed approach yields a powerful feature representation for RGB-D object recognition. Our extensive experimental evaluation on two challenging RGB-D object datasets and live video streams from Kinect shows that our learned features result in superior object recognition accuracies compared with the state-of-the-art methods.

ICRA Conference 2015 Conference Paper

Efficient RGB-D object categorization using cascaded ensembles of randomized decision trees

  • Umar Asif
  • Mohammed Bennamoun
  • Ferdous Ahmed Sohel

This paper presents an efficient framework for the categorization of objects in real-world scenes (captured with an RGB-D sensor). The proposed framework uses ensembles of randomized decision trees in a hierarchical cascaded architecture to compute consistent object-class inferences of unseen objects. Specifically, the proposed framework computes object-class probabilities at three levels of an image hierarchy (i. e. , pixel-, surfel-, and object-levels) using Random Forest classifiers. Next, these probabilities are fused together to compute a cumulative probabilistic output which is used to infer object categories. This fusion results in an improved object categorization performance compared with the state-of-the-art methods.

ICML Conference 2015 Conference Paper

How Can Deep Rectifier Networks Achieve Linear Separability and Preserve Distances?

  • Senjian An
  • Farid Boussaïd
  • Mohammed Bennamoun

This paper investigates how hidden layers of deep rectifier networks are capable of transforming two or more pattern sets to be linearly separable while preserving the distances with a guaranteed degree, and proves the universal classification power of such distance preserving rectifier networks. Through the nearly isometric nonlinear transformation in the hidden layers, the margin of the linear separating plane in the output layer and the margin of the nonlinear separating boundary in the original data space can be closely related so that the maximum margin classification in the input data space can be achieved approximately via the maximum margin linear classifiers in the output layer. The generalization performance of such distance preserving deep rectifier neural networks can be well justified by the distance-preserving properties of their hidden layers and the maximum margin property of the linear classifiers in the output layer.

IROS Conference 2014 Conference Paper

A model-free approach for the segmentation of unknown objects

  • Umar Asif
  • Mohammed Bennamoun
  • Ferdous Ahmed Sohel

We address the problem of object segmentation from depth images of highly complex indoor scenes. We propose a model-free segmentation approach, which robustly separates unknown stacked objects in real-world scenes. Our approach constructs geometrically constrained 3D clusters known as salient-regions, which are subsequently merged into high-level object hypotheses by analyzing the local geometrical characteristics (such as local shape and homogeneity) of the area of their shared boundaries. We tested our approach using depth images from live Kinect video streams and publicly available RGB-D datasets. Our approach is highly efficient and achieves superior performance compared to state-of-the-art techniques.

IROS Conference 1991 Conference Paper

Avoidance of unknown obstacles using proximity fields

  • Mohammed Bennamoun
  • Ahmad A. Masoud
  • M. A. Ramsay
  • Mohamed M. Bayoumi

Presents a novel real-time obstacle avoidance approach based on proximity sensing. The scheme is designed for the two dimensional navigation of a point object in a totally unknown environment. Navigation is performed by utilizing two proximity sensors of different settings in conjunction with a simple memory less rule for motion control. Simulation results are given for different types of obstacles. The algorithm has also been implemented on an Adept-1 arm manipulator.

v2026.09.13