Arrow Research search

Author name cluster

Zixin Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
2 author rows

Possible papers

3

NeurIPS Conference 2025 Conference Paper

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

  • Jiaben Chen
  • Zixin Wang
  • Ailing Zeng
  • Yang Fu
  • Xueyang Yu
  • Siyuan Cen
  • Julian Tanke
  • Yihang Chen

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips totaling over 500 hours of high-quality 1080P human speech videos with diverse camera shots, including close-up, half-body, and full-body views. The dataset includes detailed textual descriptions, 2D keypoints and 3D SMPL-X motion annotations, covering over 10k identities, enabling multimodal learning and evaluation. As a first attempt to showcase the value of the dataset, we present Orator, an LLM-guided multi-modal generation framework as a simple baseline, where the language model functions as a multi-faceted director, orchestrating detailed specifications for camera transitions, speaker gesticulations, and vocal modulation. This architecture enables the synthesis of coherent long-form videos through our integrated multi-modal video generation module. Extensive experiments in both pose-guided and audio-driven settings show that training on TalkCuts significantly enhances the cinematographic coherence and visual appeal of generated multi-shot speech videos. We believe TalkCuts provides a strong foundation for future work in controllable, multi-shot speech video generation and broader multimodal learning.

EAAI Journal 2025 Journal Article

Time-frequency informed stacked long short-term memory-based generative adversarial network for missing data imputation in sensor networks

  • Zixin Wang
  • Malleswari Kachireddy
  • Tarutal Ghosh Mondal
  • Wen Tang
  • Mohammad R. Jahanshahi

To monitor the health condition of civil infrastructures, the continuous acquisition of high-quality sensor data is crucial. However, in harsh environments, data can be lost due to sensor faults, data acquisition system malfunctions, or communication errors. In this study, we propose a stacked long short-term memory (LSTM)-based generative adversarial network (GAN) approach to impute missing acceleration data from faulty sensors. The stacked LSTM-based GAN model is trained by incorporating time-frequency domain information and minimizing both reconstruction and adversarial losses. The reconstruction loss is calculated based on errors in acceleration data and power spectral density (PSD). The GAN’s generator employs a stacked LSTM network to capture long-term dependencies in sequential data. The performance of our proposed approach is compared with multiple state-of-the-art methods based on GANs and variational autoencoders (VAEs). When used as baseline approaches, GANs employ a deep convolutional autoencoder (CAE) in their generators, utilizing skip connections to facilitate the flow of information from the encoder to the decoder. The proposed approach is numerically studied using a three-span continuous bridge model and the American Society of Civil Engineers (ASCE) benchmark model and experimentally validated using the physical ASCE benchmark structure and the Qatar University Grandstand Simulator (QUGS) benchmark structure. Our approach demonstrates promising performance in reconstructing missing data in both the time and frequency domains, showcasing its potential to enhance sensor fault tolerance in safety-critical systems by accurately restoring faulty sensor data.

ICML Conference 2024 Conference Paper

DNA-SE: Towards Deep Neural-Nets Assisted Semiparametric Estimation

  • Qinshuo Liu
  • Zixin Wang
  • Xi-An Li 0004
  • Xinyao Ji
  • Lei Zhang
  • Liu Lin
  • Zhonghua Liu

Semiparametric statistics play a pivotal role in a wide range of domains, including but not limited to missing data, causal inference, and transfer learning, to name a few. In many settings, semiparametric theory leads to (nearly) statistically optimal procedures that yet involve numerically solving Fredholm integral equations of the second kind. Traditional numerical methods, such as polynomial or spline approximations, are difficult to scale to multi-dimensional problems. Alternatively, statisticians may choose to approximate the original integral equations by ones with closed-form solutions, resulting in computationally more efficient, but statistically suboptimal or even incorrect procedures. To bridge this gap, we propose a novel framework by formulating the semiparametric estimation problem as a bi-level optimization problem; and then we propose a scalable algorithm called D eep N eural-Nets A ssisted S emiparametric E stimation ($\mathsf{DNA\mbox{-}SE}$) by leveraging the universal approximation property of Deep Neural-Nets (DNN) to streamline semiparametric procedures. Through extensive numerical experiments and a real data analysis, we demonstrate the numerical and statistical advantages of $\mathsf{DNA\mbox{-}SE}$ over traditional methods. To the best of our knowledge, we are the first to bring DNN into semiparametric statistics as a numerical solver of integral equations in our proposed general framework.

v2026.09.13