EUMAS Conference 2025 Demo Paper
Embedding Autonomous Agents in Resource-Constrained Robotic Platforms
- Negar Halakou
- Juan F. Gutiérrez
- Ye Sun
- Han Jiang
- Xueming Wu
- Andres Gomez
Author name cluster
Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.
EUMAS Conference 2025 Demo Paper
NeurIPS Conference 2025 Conference Paper
Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires mastering two core capabilities: video referring understanding, which captures the semantics of video regions, and video grounding, which segments object regions based on natural language descriptions. However, most existing approaches tackle these tasks in isolation, limiting progress toward unified, referentially grounded video interaction. We identify a key bottleneck in the lack of high-quality, unified video instruction data and a comprehensive benchmark for evaluating referentially grounded video chat. To address these challenges, we contribute in three core aspects: dataset, model, and benchmark. First, we introduce SAMA-239K, a large-scale dataset comprising 15K videos specifically curated to enable joint learning of video referring understanding, grounding, and multi-turn video chat. Second, we propose the SAMA model, which incorporates a versatile spatio-temporal context aggregator and a Segment Anything Model to jointly enhance fine-grained video comprehension and precise grounding capabilities. Finally, we establish SAMA-Bench, a meticulously designed benchmark consisting of 5, 067 questions from 522 videos, to comprehensively evaluate the integrated capabilities of Video LMMs in multi-turn, spatio-temporal referring understanding and grounded dialogue. Extensive experiments and benchmarking results show that SAMA not only achieves strong performance on SAMA-Bench but also sets a new state-of-the-art on general grounding benchmarks, while maintaining highly competitive performance on standard visual understanding benchmarks.
NeurIPS Conference 2024 Conference Paper
Image segmentation is a crucial vision task that groups pixels within an image into semantically meaningful segments, which is pivotal in obtaining a fine-grained understanding of real-world scenes. However, an increasing privacy concern exists regarding training large-scale image segmentation models on unauthorized private data. In this work, we exploit the concept of unlearnable examples to make images unusable to model training by generating and adding unlearnable noise into the original images. Particularly, we propose a novel Unlearnable Segmentation (UnSeg) framework to train a universal unlearnable noise generator that is capable of transforming any downstream images into their unlearnable version. The unlearnable noise generator is finetuned from the Segment Anything Model (SAM) via bilevel optimization on an interactive segmentation dataset towards minimizing the training error of a surrogate model that shares the same architecture with SAM (but trains from scratch). We empirically verify the effectiveness of UnSeg across 6 mainstream image segmentation tasks, 10 widely used datasets, and 7 different network architectures, and show that the unlearnable images can reduce the segmentation performance by a large margin. Our work provides useful insights into how to leverage foundation models in a data-efficient and computationally affordable manner to protect images against image segmentation models.
JBHI Journal 2019 Journal Article
Human skin temperature mapping provides abundant information of physiological conditions of human body, which provides supplementary or alternative indicators for disease monitoring or diagnosis. The existing models of temperature mapping or temperature field distribution of human skin are generally established by finite element method. Due to the complexity of biological systems, it is challenging to achieve high accuracy mathematical models of temperature field of human skin. The goal of this study is to establish human skin temperature three-dimensional (3-D) mapping platform by integrating optical fibers and improved genetic algorithm-back propagation (GA-BP) neural network. The proposed data-driven method is capable of acquiring entire human skin temperature 3-D mapping by simply measuring a few points on human skin. Multiple experiments were conducted to validate the proposed method on different areas of human skin in different ambient environments. In each experiment setting, the measured data and the model output data were compared. The mean absolute error in all the validation experiments is 0. 11 °C, which is lower than that in the state of the art using physical modeling for skin temperature prediction and more close to clinical accuracy. The results show that the proposed approach is accurate and reliable, which may provide a platform technology for human skin temperature mapping that can be used in both medical and scientific studies as well as home monitoring.
JBHI Journal 2014 Journal Article
This paper describes an in-vehicle nonintrusive biopotential measurement system for driver health monitoring and fatigue detection. Previous research has found that the physiological signals including eye features, electrocardiography (ECG), electroencephalography (EEG) and their secondary parameters such as heart rate and HR variability are good indicators of health state as well as driver fatigue. A conventional biopotential measurement system requires the electrodes to be in contact with human body. This not only interferes with the driver operation, but also is not feasible for long-term monitoring purpose. The driver assistance system in this paper can remotely detect the biopotential signals with no physical contact with human skin. With delicate sensor and electronic design, ECG, EEG, and eye blinking can be measured. Experiments were conducted on a high fidelity driving simulator to validate the system performance. The system was found to be able to detect the ECG/EEG signals through cloth or hair with no contact with skin. Eye blinking activities can also be detected at a distance of 10 cm. Digital signal processing algorithms were developed to decimate the signal noise and extract the physiological features. The extracted features from the vital signals were further analyzed to assess the potential criterion for alertness and drowsiness determination.