Arrow Research search

Author name cluster

Siqi Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
1 author row

Possible papers

6

NeurIPS Conference 2025 Conference Paper

PanoWan: Lifting Diffusion Video Generation Models to 360$^\circ$ with Latitude/Longitude-aware Mechanisms

  • Yifei Xia
  • Shuchen Weng
  • Siqi Yang
  • Jingqi Liu
  • Chengxuan Zhu
  • Minggui Teng
  • Zijian Jia
  • Han Jiang

Panoramic video generation enables immersive 360$^\circ$ content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-quality and diverse panoramic videos generation, due to limited dataset scale and the gap in spatial feature representations. In this paper, we introduce PanoWan to effectively lift pre-trained text-to-video models to the panoramic domain, equipped with minimal modules. PanoWan employs latitude-aware sampling to avoid latitudinal distortion, while its rotated semantic denoising and padded pixel-wise decoding ensure seamless transitions at longitude boundaries. To provide sufficient panoramic videos for learning these lifted representations, we contribute PanoVid, a high-quality panoramic video dataset with captions and diverse scenarios. Consequently, PanoWan achieves state-of-the-art performance in panoramic video generation and demonstrates robustness for zero-shot downstream tasks.

NeurIPS Conference 2025 Conference Paper

VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction-Editing Data and Long Captions

  • Ziteng Wang
  • Siqi Yang
  • Limeng Qiao
  • Lin Ma

Despite the success of Vision-Language Models (VLMs) like CLIP in aligning vision and language, their proficiency in detailed, fine-grained visual comprehension remains a key challenge. We present CLIP-IN, a novel framework that bolsters CLIP's fine-grained perception through two core innovations. Firstly, we leverage instruction-editing datasets, originally designed for image manipulation, as a unique source of hard negative image-text pairs. Coupled with a symmetric hard negative contrastive loss, this enables the model to effectively distinguish subtle visual-semantic differences. Secondly, CLIP-IN incorporates long descriptive captions, utilizing rotary positional encodings to capture rich semantic context often missed by standard CLIP. Our experiments demonstrate that CLIP-IN achieves substantial gains on the MMVP benchmark and various fine-grained visual recognition tasks, without compromising robust zero-shot performance on broader classification and retrieval tasks. Critically, integrating CLIP-IN's visual representations into Multimodal Large Language Models significantly reduces visual hallucinations and enhances reasoning abilities. This work underscores the considerable potential of synergizing targeted, instruction-based contrastive learning with comprehensive descriptive information to elevate the fine-grained understanding of VLMs.

NeurIPS Conference 2025 Conference Paper

VITRIX-UniViTAR: Unified Vision Transformer with Native Resolution

  • Limeng Qiao
  • Yiyang Gan
  • Bairui Wang
  • Jie Qin
  • Shuang Xu
  • Siqi Yang
  • Lin Ma

Conventional Vision Transformer streamlines visual modeling by employing a uniform input resolution, which underestimates the inherent variability of natural visual data and incurs a cost in spatial-contextual fidelity. While preliminary explorations have superficially investigated native resolution modeling, existing works still lack systematic training recipe from the visual representation perspective. To bridge this gap, we introduce Unified Vision Transformer with Native Resolution, i. e. UniViTAR, a family of homogeneous vision foundation models tailored for unified visual modality and native resolution scenario in the era of multimodal. Our framework first conducts architectural upgrades to the vanilla paradigm by integrating multiple advanced components. Building upon these improvements, a progressive training paradigm is introduced, which strategically combines two core mechanisms: (1) resolution curriculum learning, transitioning from fixed-resolution pretraining to native resolution tuning, thereby leveraging ViT’s inherent adaptability to variable-length sequences, and (2) visual modality adaptation via inter-batch image-video switching, which balances computational efficiency with enhanced temporal reasoning. In parallel, a hybrid training framework further synergizes sigmoid-based contrastive loss with feature distillation from a frozen teacher model, thereby accelerating early-stage convergence. Finally, trained exclusively on public accessible image-caption data, our UniViTAR family across multiple model scales from 0. 3B to 1B achieves state-of-the-art performance on a wide variety of visual-related tasks. The code and models are available here.

AAAI Conference 2021 Conference Paper

Minimizing Labeling Cost for Nuclei Instance Segmentation and Classification with Cross-domain Images and Weak Labels

  • Siqi Yang
  • Jun Zhang
  • Junzhou Huang
  • Brian C. Lovell
  • Xiao Han

Nucleus instance segmentation and classification in histopathological images is an essential prerequisite in pathology diagnosis/prognosis. However, nucleus annotations (e. g. , segmentation and labeling) require domain experts, and annotating nuclei at pixel-level is time-consuming and labor-intensive. Moreover, nuclei from different cancer types vary in shapes and appearances. These inter-cancer variations require careful annotations for specific cancer types. Therefore, to minimize the labeling cost, we propose a novel application that considers each cancer type as an individual domain and apply domain adaptation techniques to improve the segmentation/classification performance among different cancer types. Unlike the previous studies that focus on unsupervised or weakly-supervised domain adaptation independently, we would like to discover what kinds of labeling can achieve the most cost-effective domain adaptation performance in nucleus instance segmentation and classification. Specifically, we propose a unified framework that is applicable to different level annotations: no annotations, image-level, and point-level annotations. Cyclic adaptation with pseudo labels and adversarial discriminator are utilized for unsupervised domain alignment. Image-level or point-level annotations are additionally adopted to supervise the nucleus classification and refine the pseudo labels. Experiments demonstrate the effectiveness and efficacy of the proposed framework (jointly using unsupervised and weakly supervised learning) on adapting the segmentation and classification model from one cancer type to 18 other cancer types.

YNIMG Journal 2021 Journal Article

Systematically disrupted functional gradient of the cortical connectome in generalized epilepsy: Initial discovery and independent sample replication

  • Yao Meng
  • Siqi Yang
  • Huafu Chen
  • Jiao Li
  • Qiang Xu
  • Qirui Zhang
  • Guangming Lu
  • Zhiqiang Zhang

Genetic generalized epilepsy is a network disorder typically involving distributed areas identified by classical neuroanatomy. However, the finer topological relationships in terms of continuous spatial arrangement between these systems are still ambiguous. Connectome gradients provide the topological representations of human macroscale hierarchy in an abstract low-dimensional space by embedding the functional connectome into a set of axes. Leveraging connectome gradients, we systematically scrutinized abnormalities of functional connectome gradient in patients with genetic generalized epilepsy with tonic-clonic seizure (GGE-GTCS, n = 78) compared to healthy controls (HC, n = 85), and further examined the reproducibility across multiple processing configurations and in an independent validation sample (patients with GGE-GTCS, n = 28; HC, n = 31). Our findings demonstrated an extended principal gradient at different spatial scales, network-level and vertex-level, in patients with GGE-GTCS. We found consistent results across processing parameters and in validation sample. The extended principal gradient revealed the excessive functional segregation between unimodal and transmodal systems associated with duration of epilepsy and age at seizure onset in patients. Furthermore, the connectivity profile of regions with abnormal principal gradients verified the disrupted functional hierarchy revealed by gradients. Together, our findings provided a novel view of functional system hierarchy alterations, which facilitated a continuous spatial arrangement of macroscale networks, to increase our understanding of the functional connectome hierarchy in generalized epilepsy.

YNIMG Journal 2020 Journal Article

The thalamic functional gradient and its relationship to structural basis and cognitive relevance

  • Siqi Yang
  • Yao Meng
  • Jiao Li
  • Bing Li
  • Yun-Shuang Fan
  • Huafu Chen
  • Wei Liao

The human thalamus is an integrative hub richly connected with cortical networks, involving diverse cognitive functions. Emerging evidence suggests that multiscale structural and functional gradients integrate various information across modalities into an abstract representation. However, the presence of functional gradients in the thalamus and its relationship to structural properties and cognitive functions remain unknown. We estimated the functional gradients of the thalamus in two independent normal cohorts using a novel diffusion embedding analysis. We identified two main axes of the functional connectivity patterns, and examined associations with thalamic anatomy, morphology, intrinsic geometry, and specific behavioral relevance. We found that the dominant gradient indicated a lateral/medial axis across the thalamus and captured associations with anatomical nuclei and gray matter volume. The second gradient was an anterior/posterior axis and provided a behavioral characterization from lower level perception to higher level cognition. Furthermore, these two gradients strongly correlated with spatial distance, indicating the prominence of intrinsic geometry in functional hierarchies. These findings were replicated in an independent dataset. Overall, our findings suggested that macroscale gradients showed a coordination of structural and functional interactions, with hierarchical organization contributing to behavior characterization.

v2026.09.13