Arrow Research search

Author name cluster

Joo-Hwee Lim

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
1 author row

Possible papers

11

AAAI Conference 2025 System Paper

SPASCA: Social Presence and Support with Conversational Agent for Persons Living with Dementia

  • Ali Köksal
  • Jingjing Gu
  • Kotaro Hara
  • Jing Jiang
  • Joo-Hwee Lim
  • Qianli Xu

We present SPASCA - a conversational AI system that promotes psychological and cognitive well-being of persons living with dementia (PLWD). This system features an AI agent that provides social presence and support to PLWD through verbal communications, without physical presence of human caregivers. The system integrates (1) a novel dialogue model that generates dialogue items relevant to the user's experiences and lifestyle, (2) a digital avatar in the form of a talking head with the identity of a caregiver who is familiar to the demented user. We develop prototypes that adopt various interaction modalities and conversational styles and report the pros and cons of different system configurations through expert review. Our system shows the potential of conversational AI for personalized and affordable healthcare services.

NeurIPS Conference 2024 Conference Paper

MVGamba: Unify 3D Content Generation as State Space Sequence Modeling

  • Xuanyu Yi
  • Zike Wu
  • Qiuhong Shen
  • Qingshan Xu
  • Pan Zhou
  • Joo-Hwee Lim
  • Shuicheng Yan
  • Xinchao Wang

Recent 3D large reconstruction models (LRMs) can generate high-quality 3D content in sub-seconds by integrating multi-view diffusion models with scalable multi-view reconstructors. Current works further leverage 3D Gaussian Splatting as 3D representation for improved visual quality and rendering efficiency. However, we observe that existing Gaussian reconstruction models often suffer from multi-view inconsistency and blurred textures. We attribute this to the compromise of multi-view information propagation in favor of adopting powerful yet computationally intensive architectures (\eg, Transformers). To address this issue, we introduce MVGamba, a general and lightweight Gaussian reconstruction model featuring a multi-view Gaussian reconstructor based on the RNN-like State Space Model (SSM). Our Gaussian reconstructor propagates causal context containing multi-view information for cross-view self-refinement while generating a long sequence of Gaussians for fine-detail modeling with linear complexity. With off-the-shelf multi-view diffusion models integrated, MVGamba unifies 3D generation tasks from a single image, sparse images, or text prompts. Extensive experiments demonstrate that MVGamba outperforms state-of-the-art baselines in all 3D content generation scenarios with approximately only $0. 1\times$ of the model size. The codes are available at \url{https: //github. com/SkyworkAI/MVGamba}.

AAAI Conference 2023 Conference Paper

Counterfactual Dynamics Forecasting – a New Setting of Quantitative Reasoning

  • Yanzhu Liu
  • Ying Sun
  • Joo-Hwee Lim

Rethinking and introspection are important elements of human intelligence. To mimic these capabilities, counterfactual reasoning has attracted attention of AI researchers recently, which aims to forecast the alternative outcomes for hypothetical scenarios (“what-if”). However, most existing approaches focused on qualitative reasoning (e.g., casual-effect relationship). It lacks a well-defined description of the differences between counterfactuals and facts, as well as how these differences evolve over time. This paper defines a new problem formulation - counterfactual dynamics forecasting - which is described in middle-level abstraction under the structural causal models (SCM) framework and derived as ordinary differential equations (ODEs) as low-level quantitative computation. Based on it, we propose a method to infer counterfactual dynamics considering the factual dynamics as demonstration. Moreover, the evolution of differences between facts and counterfactuals are modelled by an explicit temporal component. The experimental results on two dynamical systems demonstrate the effectiveness of the proposed method.

NeurIPS Conference 2021 Conference Paper

Predicting Event Memorability from Contextual Visual Semantics

  • Qianli Xu
  • Fen Fang
  • Ana Molino
  • Vigneshwaran Subbaraju
  • Joo-Hwee Lim

Episodic event memory is a key component of human cognition. Predicting event memorability, i. e. , to what extent an event is recalled, is a tough challenge in memory research and has profound implications for artificial intelligence. In this study, we investigate factors that affect event memorability according to a cued recall process. Specifically, we explore whether event memorability is contingent on the event context, as well as the intrinsic visual attributes of image cues. We design a novel experiment protocol and conduct a large-scale experiment with 47 elder subjects over 3 months. Subjects’ memory of life events is tested in a cued recall process. Using advanced visual analytics methods, we build a first-of-its-kind event memorability dataset (called R3) with rich information about event context and visual semantic features. Furthermore, we propose a contextual event memory network (CEMNet) that tackles multi-modal input to predict item-wise event memorability, which outperforms competitive benchmarks. The findings inform deeper understanding of episodic event memory, and open up a new avenue for prediction of human episodic memory. Source code is available at https: //github. com/ffzzy840304/Predicting-Event-Memorability.

AAAI Conference 2021 System Paper

TAILOR: Teaching with Active and Incremental Learning for Object Registration

  • Qianli Xu
  • Nicolas Gauthier
  • Wenyu Liang
  • Fen Fang
  • Hui Li Tan
  • Ying Sun
  • Yan Wu
  • Liyuan Li

When deploying a robot to a new task, one often has to train it to detect novel objects, which is time-consuming and laborintensive. We present TAILOR - a method and system for object registration with active and incremental learning. When instructed by a human teacher to register an object, TAILOR is able to automatically select viewpoints to capture informative images by actively exploring viewpoints, and employs a fast incremental learning algorithm to learn new objects without potential forgetting of previously learned objects. We demonstrate the effectiveness of our method with a KUKA robot to learn novel objects used in a real-world gearbox assembly task through natural interactions.

AAAI Conference 2019 Conference Paper

Singe Image Rain Removal with Unpaired Information: A Differentiable Programming Perspective

  • Hongyuan Zhu
  • Xi Peng
  • Joey Tianyi Zhou
  • Songfan Yang
  • Vijay Chanderasekh
  • Liyuan Li
  • Joo-Hwee Lim

Single image rain-streak removal is an extremely challenging problem due to the presence of non-uniform rain densities in images. Previous works solve this problem using various hand-designed priors or by explicitly mapping synthetic rain to paired clean image in a supervised way. In practice, however, the pre-defined priors are easily violated and the paired training data are hard to collect. To overcome these limitations, in this work, we propose RainRemoval-GAN (RR- GAN), the first end-to-end adversarial model that generates realistic rain-free images using only unpaired supervision. Our approach alleviates the paired training constraints by introducing a physical-model which explicitly learns a recovered images and corresponding rain-streaks from the differentiable programming perspective. The proposed network consists of a novel multiscale attention memory generator and a novel multiscale deeply supervised discriminator. The multiscale attention memory generator uses a memory with attention mechanism to capture the latent rain streaks context at different stages to recover the clean images. The deeply supervised multiscale discriminator imposes constraints at the recovered output in terms of local details and global appearance to the clean image set. Together with the learned rainstreaks, a reconstruction constraint is employed to ensure the appearance consistent with the input image. Experimental results on public benchmark demonstrates our promising performance compared with nine state-of-the-art methods in terms of PSNR, SSIM, visual qualities and running time.

IJCAI Conference 2018 Conference Paper

DehazeGAN: When Image Dehazing Meets Differential Programming

  • Hongyuan Zhu
  • Xi Peng
  • Vijay Chandrasekhar
  • Liyuan Li
  • Joo-Hwee Lim

Single image dehazing has been a classic topic in computer vision for years. Motivated by the atmospheric scattering model, the key to satisfactory single image dehazing relies on an estimation of two physical parameters, i. e. , the global atmospheric light and the transmission coefficient. Most existing methods employ a two-step pipeline to estimate these two parameters with heuristics which accumulate errors and compromise dehazing quality. Inspired by differentiable programming, we re-formulate the atmospheric scattering model into a novel generative adversarial network (DehazeGAN). Such a reformulation and adversarial learning allow the two parameters to be learned simultaneously and automatically from data by optimizing the final dehazing performance so that clean images with faithful color and structures are directly produced. Moreover, our reformulation also greatly improves the GAN’s interpretability and quality for single image dehazing. To the best of our knowledge, our method is one of the first works to explore the connection among generative adversarial models, image dehazing, and differentiable programming, which advance the theories and application of these areas. Extensive experiments on synthetic and realistic data show that our method outperforms state-of-the-art methods in terms of PSNR, SSIM, and subjective visual quality.

AAAI Conference 2017 Conference Paper

Active Video Summarization: Customized Summaries via On-line Interaction with the User

  • Ana Garcia del Molino
  • Xavier Boix
  • Joo-Hwee Lim
  • Ah-Hwee Tan

To facilitate the browsing of long videos, automatic video summarization provides an excerpt that represents its content. In the case of egocentric and consumer videos, due to their personal nature, adapting the summary to specific user’s preferences is desirable. Current approaches to customizable video summarization obtain the user’s preferences prior to the summarization process. As a result, the user needs to manually modify the summary to further meet the preferences. In this paper, we introduce Active Video Summarization (AVS), an interactive approach to gather the user’s preferences while creating the summary. AVS asks questions about the summary to update it on-line until the user is satisfied. To minimize the interaction, the best segment to inquire next is inferred from the previous feedback. We evaluate AVS in the commonly used UTEgo dataset. We also introduce a new dataset for customized video summarization (CSumm) recorded with a Google Glass. The results show that AVS achieves an excellent compromise between usability and quality. In 41% of the videos, AVS is considered the best over all tested baselines, including summaries manually generated. Also, when looking for specific events in the video, AVS provides an average level of satisfaction higher than those of all other baselines after only six questions to the user.

NeurIPS Conference 2013 Conference Paper

Top-Down Regularization of Deep Belief Networks

  • Hanlin Goh
  • Nicolas Thome
  • Matthieu Cord
  • Joo-Hwee Lim

Designing a principled and effective algorithm for learning deep architectures is a challenging problem. The current approach involves two training phases: a fully unsupervised learning followed by a strongly discriminative optimization. We suggest a deep learning strategy that bridges the gap between the two phases, resulting in a three-phase learning procedure. We propose to implement the scheme using a method to regularize deep belief networks with top-down information. The network is constructed from building blocks of restricted Boltzmann machines learned by combining bottom-up and top-down sampled signals. A global optimization procedure that merges samples from a forward bottom-up pass and a top-down pass is used. Experiments on the MNIST dataset show improvements over the existing algorithms for deep belief networks. Object recognition results on the Caltech-101 dataset also yield competitive results.

IJCAI Conference 2003 Conference Paper

Learning Consumer Photo Categories for Semantic Retrieval

  • Joo-Hwee Lim
  • Jesse S. Jin

In this paper, wo develop a computational learning framework to build a hierarchy of 11 consumer photo categories for semantic retrieval. Two levels of visual semantics are learned for image content and image category statistically. We evaluate the average precisions at top retrieved photos on 2400 heterogeneous consumer photos with very good result.

AAAI Conference 1992 Conference Paper

A Framework for Integrating Fault Diagnosis and Incremental Knowledge Acquisition in Connectionist Expert Systems

  • Joo-Hwee Lim

In this paper, we propose a framework for integrating fault diagnosis and incremental knowledge acquisition in connectionist expert systems. A new case solved by the Diagnostic Function is formulated as a new example for the Learning Function to learn incrementally. The Diagnostic Function is composed of a neural networks-based Example Module and a symbolic-based Rule Module. While the Example Module is always first invoked to provide the short-cut solution, the Rule Module provides extensive coverage of cases to handle odd cases when Example Module fails. Two applications based on the proposed framework will also be briefly mentioned.

v2026.09.13