Arrow Research search

Author name cluster

Yiran Shen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
1 author row

Possible papers

5

AAAI Conference 2026 Conference Paper

CreBench: Human-Aligned Creativity Evaluation from Idea to Process to Product

  • Kaiwen Xue
  • Chenglong Li
  • Zhonghong Ou
  • Guoxin Zhang
  • Kaoyan Lu
  • Shuai Lyu
  • Yifan Zhu
  • Ping Zong

Human-defined creativity is highly abstract, posing a challenge for multimodal large language models (MLLMs) to comprehend and assess creativity that aligns with human judgments. The absence of an existing benchmark further exacerbates this dilemma. To this end, we propose CreBench, which consists of two key components: 1) an evaluation benchmark covering the multiple dimensions from creative idea to process to products; 2) CreMIT (Creativity Multimodal Instruction Tuning dataset), a multimodal creativity evaluation dataset, consisting of 2.2K diverse-sourced multimodal data, 79.2K human feedbacks and 4.7M multityped instructions. Specifically, to ensure MLLMs can handle diverse creativity-related queries, we prompt GPT to refine the human feedback to activate stronger creativity assessment capabilities. CreBench serves as a foundation for building MLLMs that understand human-aligned creativity. Based on the CreBench, we fine-tune open-source general MLLMs, resulting in CreExpert, a multimodal creativity evaluation expert model. Extensive experiments demonstrate that the proposed CreExpert models achieve significantly better alignment with human creativity evaluation compared to state-ofthe-art MLLMs, including the most advanced GPT-4V and Gemini-Pro-Vision.

AAAI Conference 2026 Conference Paper

Ev-iCRF: Self-supervised Event-guided iCRF Estimation for HDR Image Reconstruction

  • Xucheng Guo
  • Bing Li
  • Lin Wang
  • Yiran Shen

In this paper, we present Ev-iCRF, a novel self-supervised pipeline for high dynamic range (HDR) image reconstruction from a single-exposure low dynamic range (LDR) image, guided by asynchronous event streams generated by a bio-inspired event camera. The highlight of Ev-iCRF lies in its formulation of the inverse camera response function (iCRF) based on Event-LDR Correspondence. By leveraging the HDR properties of event data, the method enables direct iCRF estimation, offering a new perspective for event-guided HDR imaging. The pipeline is trained in a self-supervised manner using formulation-driven iCRF estimation loss and refinement loss, without the need for synchronized HDR supervision. Ev-iCRF adopts a two-stage coarse-to-fine reconstruction pipeline, allowing effective fusion of features from both LDR image and event data. The event information is used to optimize the iCRF, enabling accurate HDR reconstruction from LDR inputs. We evaluate Ev-iCRF on real-world datasets, and results show that it outperforms state-of-the-art methods in HDR reconstruction accuracy. Moreover, the reconstructed images demonstrate improved texture fidelity and structural detail.

AAAI Conference 2024 Conference Paper

E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning

  • Qiang Qu
  • Yiran Shen
  • Xiaoming Chen
  • Yuk Ying Chung
  • Tongliang Liu

The bio-inspired event cameras or dynamic vision sensors are capable of asynchronously capturing per-pixel brightness changes (called event-streams) in high temporal resolution and high dynamic range. However, the non-structural spatial-temporal event-streams make it challenging for providing intuitive visualization with rich semantic information for human vision. It calls for events-to-video (E2V) solutions which take event-streams as input and generate high quality video frames for intuitive visualization. However, current solutions are predominantly data-driven without considering the prior knowledge of the underlying statistics relating event-streams and video frames. It highly relies on the non-linearity and generalization capability of the deep neural networks, thus, is struggling on reconstructing detailed textures when the scenes are complex. In this work, we propose E2HQV, a novel E2V paradigm designed to produce high-quality video frames from events. This approach leverages a model-aided deep learning framework, underpinned by a theory-inspired E2V model, which is meticulously derived from the fundamental imaging principles of event cameras. To deal with the issue of state-reset in the recurrent components of E2HQV, we also design a temporal shift embedding module to further improve the quality of the video frames. Comprehensive evaluations on the real world event camera datasets validate our approach, with E2HQV, notably outperforming state-of-the-art approaches, e.g., surpassing the second best by over 40% for some evaluation metrics.

NeurIPS Conference 2023 Conference Paper

EV-Eye: Rethinking High-frequency Eye Tracking through the Lenses of Event Cameras

  • Guangrong Zhao
  • Yurun Yang
  • Jingwei Liu
  • Ning Chen
  • Yiran Shen
  • Hongkai Wen
  • Guohao Lan

In this paper, we present EV-Eye, a first-of-its-kind large scale multimodal eye tracking dataset aimed at inspiring research on high-frequency eye/gaze tracking. EV-Eye utilizes an emerging bio-inspired event camera to capture independent pixel-level intensity changes induced by eye movements, achieving sub-microsecond latency. Our dataset was curated over a two-week period and collected from 48 participants encompassing diverse genders and age groups. It comprises over 1. 5 million near-eye grayscale images and 2. 7 billion event samples generated by two DAVIS346 event cameras. Additionally, the dataset contains 675 thousands scene images and 2. 7 million gaze references captured by Tobii Pro Glasses 3 eye tracker for cross-modality validation. Compared with existing event-based high-frequency eye tracking datasets, our dataset is significantly larger in size, and the gaze references involve more natural eye movement patterns, i. e. , fixation, saccade and smooth pursuit. Alongside the event data, we also present a hybrid eye tracking method as benchmark, which leverages both the near-eye grayscale images and event data for robust and high-frequency eye tracking. We show that our method achieves higher accuracy for both pupil and gaze estimation tasks compared to the existing solution.

EAAI Journal 2023 Journal Article

VasLine: Realize online detection and augmented NIR using deep learning

  • Zhongxin Chen
  • Yiran Shen
  • Binbin Chen
  • Jun Zhou
  • Panling Huang
  • Hengchang Zang
  • Yongxia Guan

The chemical information acquisition capabilities of near infrared (NIR) have attracted great interests of investigators on exploring its potential as process analytical technologies (PAT) on the traditional Chinese medicine (TCM). However, due to the operation environment of TCM is complex and noisy, the accurate online composition detection and analysis is challenging and remains unsolved. Therefore, in this paper, we propose a new online TCM composition analysis platform for high quality online spectral data acquisition. We design, VasLine, a generative framework based on deep learning to estimate the multiple critical quality attributes (CQAs). To demonstrate the feasibility of the system, the platform was applied to the extraction process of Xiao’er Xiaoji Zhike Oral Liquid (XXZOL) and the framework was trained to estimate the content of the 7 CQAs. In the evaluation, we carried out 16 batches extraction experiments and collected extensive online NIR data, offline NIR data and high-performance liquid chromatography (HPLC) data for TCM composition analysis. The results show that the estimation accuracy of VasLine outperforms the state-of-the-art regression approaches significantly for all the 7 different CQAs, e. g. , the R2 metrics of VasLine for the CQAs are all higher than 0. 95. This study aims to propose a novel generative framework based on a deep learning model, and the framework is applied in the self-developed platform to estimate CQAs with high-quality generative data from noisy online NIR accurately. The experiments for online TCM composition analysis show that VasLine is a new solution for the quality improvement of the actual pharmaceutical industry.

v2026.09.13