Arrow Research search

Author name cluster

Yichen Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

22 papers
2 author rows

Possible papers

22

EAAI Journal 2026 Journal Article

A global and local agent-based curriculum reinforcement learning approach for multi-end-effector robotic arm manipulation

  • Yichen Wang
  • Shuai Zheng
  • Ze Yang
  • Jingmin Guo
  • Zitong Yang
  • Jun Hong

Reinforcement learning is widely applied in robotic arm manipulation tasks. However, most of these tasks focus on single and simple end effector. When facing heavy robotic arm hoisting tasks, which are usually manipulated by robotic arms with multi-end-effectors and more degrees of freedom, the single-agent-based reinforcement learning method performs relatively ineffective. In this paper, we propose a multi-agent reinforcement learning approach for hoisting tasks manipulated by robotic arm with multi-end-effectors. The method decomposes the robotic arm into global and local agents based on the degrees of freedom, with one agent controlling global and rough movement, and the other controlling local and fine movement. In this way, the multi-end-effectors’ spatial trajectory can be accurately manipulated. Moreover, in the training process, a four levels curriculum learning strategy is introduced, in which different reward functions are designed respectively, to make the training efficiency and effectiveness. We develop a Unity engine environment-based simulation and perform several comparison experiments. The results demonstrate that the proposed approach outperforms conventional single-agent-based methods.

NeurIPS Conference 2025 Conference Paper

AdvEDM: Fine-grained Adversarial Attack against VLM-based Embodied Agents

  • Yichen Wang
  • Hangtao Zhang
  • Hewen Pan
  • Ziqi Zhou
  • Xianlong Wang
  • Peijin Guo
  • Lulu Xue
  • Shengshan Hu

Vision-Language Models (VLMs), with their strong reasoning and planning capabilities, are widely used in embodied decision-making (EDM) tasks in embodied agents, such as autonomous driving and robotic manipulation. Recent research has increasingly explored adversarial attacks on VLMs to reveal their vulnerabilities. However, these attacks either rely on overly strong assumptions, requiring full knowledge of the victim VLM, which is impractical for attacking VLM-based agents, or exhibit limited effectiveness. The latter stems from disrupting most semantic information in the image, which leads to a misalignment between the perception and the task context defined by system prompts. This inconsistency interrupts the VLM's reasoning process, resulting in invalid outputs that fail to affect interactions in the physical world. To this end, we propose a fine-grained adversarial attack framework, AdvEDM, which modifies the VLM's perception of only a few key objects while preserving the semantics of the remaining regions. This attack effectively reduces conflicts with the task context, making VLMs output valid but incorrect decisions and affecting the actions of agents, thus posing a more substantial safety threat in the physical world. We design two variants of based on this framework, AdvEDM-R and AdvEDM-A, which respectively remove the semantics of a specific object from the image and add the semantics of a new object into the image. The experimental results in both general scenarios and EDM tasks demonstrate fine-grained control and excellent attack performance.

ICLR Conference 2025 Conference Paper

BadRobot: Jailbreaking Embodied LLM Agents in the Physical World

  • Hangtao Zhang
  • Chenyu Zhu
  • Xianlong Wang 0001
  • Ziqi Zhou 0001
  • Changgan Yin
  • Minghui Li
  • Lulu Xue
  • Yichen Wang

Embodied AI represents systems where AI is integrated into physical entities. Multimodal Large Language Model (LLM), which exhibits powerful language understanding abilities, has been extensively employed in embodied AI by facilitating sophisticated task planning. However, a critical safety issue remains overlooked: could these embodied LLMs perpetrate harmful behaviors? In response, we introduce BadRobot, the first attack paradigm designed to jailbreak robotic manipulation, making embodied LLMs violate safety and ethical constraints through typical voice-based user-system interactions. Specifically, three vulnerabilities are exploited to achieve this type of attack: (i) manipulation of LLMs within robotic systems, (ii) misalignment between linguistic outputs and physical actions, and (iii) unintentional hazardous behaviors caused by world knowledge's flaws. Furthermore, we construct a benchmark of various malicious physical action queries to evaluate BadRobot's attack performance. Based on this benchmark, extensive experiments against existing prominent embodied LLM frameworks (e.g., Voxposer, Code as Policies, and ProgPrompt) demonstrate the effectiveness of our BadRobot. We emphasize that addressing this emerging vulnerability is crucial for the secure deployment of LLMs in robotics. Warning: This paper contains harmful AI-generated language and aggressive actions.

AAAI Conference 2025 Conference Paper

Breaking Barriers in Physical-World Adversarial Examples: Improving Robustness and Transferability via Robust Feature

  • Yichen Wang
  • Yuxuan Chou
  • Ziqi Zhou
  • Hangtao Zhang
  • Wei Wan
  • Shengshan Hu
  • Minghui Li

As deep neural networks (DNNs) are widely applied in the physical world, many researches are focusing on physical-world adversarial examples (PAEs), which introduce perturbations to inputs and cause the model's incorrect outputs. However, existing PAEs face two challenges: unsatisfactory attack performance (i.e., poor transferability and insufficient robustness to environment conditions), and difficulty in balancing attack effectiveness with stealthiness, where better attack effectiveness often makes PAEs more perceptible. In this paper, we explore a novel perturbation-based method to overcome the challenges. For the first challenge, we introduce a strategy Deceptive RF injection based on robust features (RFs) that are predictive, robust to perturbations, and consistent across different models. Specifically, it improves the transferability and robustness of PAEs by covering RFs of other classes onto the predictive features in clean images. For the second challenge, we introduce another strategy Adversarial Semantic Pattern Minimization, which removes most perturbations and retains only essential adversarial patterns in AEs. Based on the two strategies, we design our method Robust Feature Coverage Attack (RFCoA), comprising Robust Feature Disentanglement and Adversarial Feature Fusion. In the first stage, we extract target class RFs in feature space. In the second stage, we use attention-based feature fusion to overlay these RFs onto predictive features of clean images and remove unnecessary perturbations. Experiments show our method's superior transferability, robustness, and stealthiness compared to existing state-of-the-art methods. Additionally, our method's effectiveness can extend to Large Vision-Language Models (LVLMs), indicating its potential applicability to more complex tasks.

ICML Conference 2025 Conference Paper

Feature Learning beyond the Lazy-Rich Dichotomy: Insights from Representational Geometry

  • Chi-Ning Chou
  • Hang Le 0005
  • Yichen Wang
  • SueYeon Chung

Integrating task-relevant information into neural representations is a fundamental ability of both biological and artificial intelligence systems. Recent theories have categorized learning into two regimes: the rich regime, where neural networks actively learn task-relevant features, and the lazy regime, where networks behave like random feature models. Yet this simple lazy–rich dichotomy overlooks a diverse underlying taxonomy of feature learning, shaped by differences in learning algorithms, network architectures, and data properties. To address this gap, we introduce an analysis framework to study feature learning via the geometry of neural representations. Rather than inspecting individual learned features, we characterize how task-relevant representational manifolds evolve throughout the learning process. We show, in both theoretical and empirical settings, that as networks learn features, task-relevant manifolds untangle, with changes in manifold geometry revealing distinct learning stages and strategies beyond the lazy–rich dichotomy. This framework provides novel insights into feature learning across neuroscience and machine learning, shedding light on structural inductive biases in neural circuits and the mechanisms underlying out-of-distribution generalization.

ICRA Conference 2025 Conference Paper

Kinodynamic Model Predictive Control for Energy Efficient Locomotion of Legged Robots with Parallel Elasticity

  • Yulun Zhuang
  • Yichen Wang
  • Yanran Ding

In this paper, we introduce a kinodynamic model predictive control (MPC) framework that exploits unidirectional parallel springs (UPS) to improve the energy efficiency of dynamic legged robots. The proposed method employs a hierarchical control structure, where the solution of MPC with simplified dynamic models is used to warm-start the kinody-namic MPC, which accounts for nonlinear centroidal dynamics and kinematic constraints. The proposed approach enables energy efficient dynamic hopping on legged robots by using UPS to reduce peak motor torques and energy consumption during stance phases. Simulation results demonstrated a 38. 8% reduction in the cost of transport (CoT) for a monoped robot equipped with UPS during high-speed hopping. Additionally, preliminary hardware experiments show a 14. 8% reduction in energy consumption.

YNIMG Journal 2025 Journal Article

Multimodal integration of plasma biomarkers, MRI, and genetic risk to predict cerebral amyloid burden in Alzheimer’s disease

  • Yichen Wang
  • Hao-Jie Chen
  • Yuxin Cheng
  • YaoXin Xie
  • Yuyan Cheng
  • Shiyun Zhao
  • Yidong Jiang
  • Tianyu Bai

Alzheimer’s disease (AD), the most prevalent neurodegenerative disorder, is marked by the accumulation of amyloid-β (Aβ) plaques. Although cerebral Aβ positron emission tomography (Aβ-PET) remains the gold standard for assessing cerebral Aβ burden, its clinical utility is hindered by cost, radiation exposure, and limited availability. Plasma biomarkers have emerged as promising, non‑invasive indicators of Aβ pathology, yet they do not incorporate individual genetic risk or neuroanatomical context. To address this gap, we developed a multimodal machine‑learning framework that integrates plasma biomarkers, MRI‑derived brain structural features (regional volumes, cortical thickness, cortical area and structural connectivity), and genetic risk profiles to predict cerebral Aβ burden. This approach was evaluated in 150 participants from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) and 101 participants from a domestic Chinese Sino Longitudinal Study of Cognitive Decline (SILCODE). Incorporating multimodal features substantially improved predictive performance: the baseline model using plasma and clinical variables alone achieved an R2 of 0. 56, whereas integrating neuroimaging and genetic information increased accuracy (R2 = 0. 63 with apolipoprotein E genotypes and R2 = 0. 64 with polygenic risk scores). Furthermore, a multiclass classifier trained on the same multimodal features achieved robust discrimination of cognitive status, with area‑under‑the‑curve values of 0. 87 for normal controls, 0. 76 for mild cognitive impairment, and 0. 95 for AD dementia. These findings highlight the value of combining plasma, imaging, and genetic data to non-invasively estimate cerebral Aβ burden, offering a potential alternative to PET imaging for early AD risk assessment.

IROS Conference 2025 Conference Paper

New Network Protocol for Supermedia-Enhanced Telerobotics

  • Xinyu Liu 0013
  • Zekun Song
  • Hongli Huang
  • Yuxuan Xue
  • Yichen Wang
  • Vellaisamy A. L. Roy
  • Ning Xi 0001

The growing complexity of robotic teleoperation systems necessitates the integration of multiple feedback modalities, including video, audio, force, tactile, and temperature feedback. The concept of supermedia is utilized to describe the aggregation of these feedback streams. By integrating multiple media forms, supermedia can offer a more comprehensive interactive experience for robot teleoperation systems. However, existing transmission protocols struggle to maintain synchronization among these diverse feedback streams, particularly in demanding network environments. In this paper, we present the Tele-Robotic Control Protocol (TRCP), a novel network transmission protocol specifically designed for supermedia-enhanced robotic teleoperation systems. TRCP incorporates an event reference mechanism that coordinates multiple feedback streams based on robot state rather than traditional time-based sampling. It also employs multi-queue management for the independent handling of different feedback types and integrates an adaptive adjustment mechanism that optimizes transmission parameters in response to real-time network conditions. The effectiveness of TRCP is demonstrated through a cross-continental teleoperation experiment between the University of Glasgow and the University of Hong Kong. TRCP achieves superior feedback synchronization and real-time responsiveness, significantly enhancing both task success rates and operator performance.

ICLR Conference 2025 Conference Paper

Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and Modalities

  • Yichen Wang
  • Yiyi Zhang
  • Xinhao Hu
  • Li Niu 0002
  • Jianfu Zhang 0003
  • Yasushi Makihara
  • Yasushi Yagi
  • Pai Peng

Reconstructing pedestrian motion from dynamic sensors, with a focus on pedestrian intention, is crucial for advancing autonomous driving safety. However, this task is challenging due to data limitations arising from technical complexities, safety, and cost concerns. We introduce the Pedestrian Motion Reconstruction (PMR) dataset, which focuses on pedestrian intention to reconstruct behavior using multiple perspectives and modalities. PMR is developed from a mixed reality platform that combines real-world realism with the extensive, accurate labels of simulations, thereby reducing costs and risks. It captures the intricate dynamics of pedestrian interactions with objects and vehicles, using different modalities for a comprehensive understanding of human-vehicle interaction. Analyses show that PMR can naturally exhibit pedestrian intent and simulate extreme cases. PMR features a vast collection of data from 54 subjects interacting across 12 urban settings with 7 objects, encompassing 12,138 sequences with diverse weather conditions and vehicle speeds. This data provides a rich foundation for modeling pedestrian intent through multi-view and multi-modal insights. We also conduct comprehensive benchmark assessments across different modalities to thoroughly evaluate pedestrian motion reconstruction methods.

NeurIPS Conference 2025 Conference Paper

The $\varphi$ Curve: The Shape of Generalization through the Lens of Norm-based Capacity Control

  • Yichen Wang
  • Yudong Chen
  • Lorenzo Rosasco
  • Fanghui Liu

Understanding how the test risk scales with model complexity is a central question in machine learning. Classical theory is challenged by the learning curves observed for large over-parametrized deep networks. Capacity measures based on parameter count typically fail to account for these empirical observations. To tackle this challenge, we consider norm-based capacity measures and develop our study for random features based estimators, widely used as simplified theoretical models for more complex networks. In this context, we provide a precise characterization of how the estimator’s norm concentrates and how it governs the associated test error. Our results show that the predicted learning curve admits a phase transition from under- to over-parameterization, but no double descent behavior. This confirms that more classical U-shaped behavior is recovered considering appropriate capacity measures based on models norms rather than size. From a technical point of view, we leverage deterministic equivalence as the key tool and further develop new deterministic quantities which are of independent interest.

NeurIPS Conference 2024 Conference Paper

Concentrate Attention: Towards Domain-Generalizable Prompt Optimization for Language Models

  • Chengzhengxu Li
  • Xiaoming Liu
  • Zhaohan Zhang
  • Yichen Wang
  • Chen Liu
  • Yu Lan
  • Chao Shen

Recent advances in prompt optimization have notably enhanced the performance of pre-trained language models (PLMs) on downstream tasks. However, the potential of optimized prompts on domain generalization has been under-explored. To explore the nature of prompt generalization on unknown domains, we conduct pilot experiments and find that (i) Prompts gaining more attention weight from PLMs’ deep layers are more generalizable and (ii) Prompts with more stable attention distributions in PLMs’ deep layers are more generalizable. Thus, we offer a fresh objective towards domain-generalizable prompts optimization named ''Concentration'', which represents the ''lookback'' attention from the current decoding token to the prompt tokens, to increase the attention strength on prompts and reduce the fluctuation of attention distribution. We adapt this new objective to popular soft prompt and hard prompt optimization methods, respectively. Extensive experiments demonstrate that our idea improves comparison prompt optimization methods by 1. 42% for soft prompt generalization and 2. 16% for hard prompt generalization in accuracy on the multi-source domain generalization setting, while maintaining satisfying in-domain performance. The promising results validate the effectiveness of our proposed prompt optimization objective and provide key insights into domain-generalizable prompts.

IJCAI Conference 2024 Conference Paper

DarkFed: A Data-Free Backdoor Attack in Federated Learning

  • Minghui Li
  • Wei Wan
  • Yuxuan Ning
  • Shengshan Hu
  • Lulu Xue
  • Leo Yu Zhang
  • Yichen Wang

Federated learning (FL) has been demonstrated to be susceptible to backdoor attacks. However, existing academic studies on FL backdoor attacks rely on a high proportion of real clients with main task-related data, which is impractical. In the context of real-world industrial scenarios, even the simplest defense suffices to defend against the state-of-the-art attack, 3DFed. A practical FL backdoor attack remains in a nascent stage of development. To bridge this gap, we present DarkFed. Initially, we emulate a series of fake clients, thereby achieving the attacker proportion typical of academic research scenarios. Given that these emulated fake clients lack genuine training data, we further propose a data-free approach to backdoor FL. Specifically, we delve into the feasibility of injecting a backdoor using a shadow dataset. Our exploration reveals that impressive attack performance can be achieved, even when there is a substantial gap between the shadow dataset and the main task dataset. This holds true even when employing synthetic data devoid of any semantic information as the shadow dataset. Subsequently, we strategically construct a series of covert backdoor updates in an optimized manner, mimicking the properties of benign updates, to evade detection by defenses. A substantial body of empirical evidence validates the tangible effectiveness of DarkFed.

IJCAI Conference 2024 Conference Paper

Detector Collapse: Backdooring Object Detection to Catastrophic Overload or Blindness in the Physical World

  • Hangtao Zhang
  • Shengshan Hu
  • Yichen Wang
  • Leo Yu Zhang
  • Ziqi Zhou
  • Xianlong Wang
  • Yanjun Zhang
  • Chao Chen

Object detection tasks, crucial in safety-critical systems like autonomous driving, focus on pinpointing object locations. These detectors are known to be susceptible to backdoor attacks. However, existing backdoor techniques have primarily been adapted from classification tasks, overlooking deeper vulnerabilities specific to object detection. This paper is dedicated to bridging this gap by introducing Detector Collapse (DC), a brand-new backdoor attack paradigm tailored for object detection. DC is designed to instantly incapacitate detectors (i. e. , severely impairing detector's performance and culminating in a denial-of-service). To this end, we develop two innovative attack schemes: Sponge for triggering widespread misidentifications and Blinding for rendering objects invisible. Remarkably, we introduce a novel poisoning strategy exploiting natural objects, enabling DC to act as a practical backdoor in real-world environments. Our experiments on different detectors across several benchmarks show a significant improvement (~10%-60% absolute and ~2-7x relative) in attack efficacy over state-of-the-art attacks.

AAAI Conference 2024 Conference Paper

Dialogue for Prompting: A Policy-Gradient-Based Discrete Prompt Generation for Few-Shot Learning

  • Chengzhengxu Li
  • Xiaoming Liu
  • Yichen Wang
  • Duyi Li
  • Yu Lan
  • Chao Shen

Prompt-based pre-trained language models (PLMs) paradigm has succeeded substantially in few-shot natural language processing (NLP) tasks. However, prior discrete prompt optimization methods require expert knowledge to design the base prompt set and identify high-quality prompts, which is costly, inefficient, and subjective. Meanwhile, existing continuous prompt optimization methods improve the performance by learning the ideal prompts through the gradient information of PLMs, whose high computational cost, and low readability and generalizability are often concerning. To address the research gap, we propose a Dialogue-comprised Policy-gradient-based Discrete Prompt Optimization (DP_2O) method. We first design a multi-round dialogue alignment strategy for readability prompt set generation based on GPT-4. Furthermore, we propose an efficient prompt screening metric to identify high-quality prompts with linear complexity. Finally, we construct a reinforcement learning (RL) framework based on policy gradients to match the prompts to inputs optimally. By training a policy network with only 0.62M parameters on the tasks in the few-shot setting, DP_2O outperforms the state-of-the-art (SOTA) method by 1.52% in accuracy on average on four open-source datasets. Moreover, subsequent experiments also demonstrate that DP_2O has good universality, robustness and generalization ability.

IROS Conference 2024 Conference Paper

Enhancing Exploratory Capability of Visual Navigation Using Uncertainty of Implicit Scene Representation

  • Yichen Wang
  • Qiming Liu 0001
  • Zhe Liu 0022
  • Hesheng Wang 0001

In the context of visual navigation in unknown scenes, both “exploration” and “exploitation” are equally crucial. Robots must first establish environmental cognition through exploration and then utilize the cognitive information to accomplish target searches. However, most existing methods for image-goal navigation prioritize target search over the generation of exploratory behavior. To address this, we propose the Navigation with Uncertainty-driven Exploration (NUE) pipeline, which uses an implicit and compact scene representation, NeRF, as a cognitive structure. We estimate the uncertainty of NeRF and augment the exploratory ability by the uncertainty to in turn facilitate the construction of implicit representation. Simultaneously, we extract memory information from NeRF to enhance the robot’s reasoning ability for determining the location of the target. Ultimately, we seamlessly combine the two generated abilities to produce navigational actions. Our pipeline is end-to-end, with the environmental cognitive structure being constructed online. Extensive experimental results on image-goal navigation demonstrate the capability of our pipeline to enhance exploratory behaviors, while also enabling a natural transition from the exploration to exploitation phase. This enables our model to outperform existing memory-based cognitive navigation structures in terms of navigation performance. Project page: https://github.com/IRMVLab/NUE-NeRF-nav

ECAI Conference 2023 Conference Paper

GraphSA: Smart Contract Vulnerability Detection Combining Graph Neural Networks and Static Analysis

  • Long He
  • Xiangfu Zhao
  • Yichen Wang
  • Jiahui Yang
  • Xuelei Sun

Security incidents in smart contracts still occur frequently, as the underlying code is often vulnerable to attacks. However, traditional methods to detect vulnerabilities in smart contracts are limited by certain rigid rules, reducing accuracy and scalability. In this work, we propose GraphSA, which combines Graph neural networks (GNNs) and Static Analysis for smart contract vulnerability detection. First, we present the contract tree, which is obtained by converting the control flow graph (CFG) of a smart contract. Each node in the tree represents a crucial operation code (opcode) block, and each edge represents the control flow (execution order) between code blocks. Then, we propose an extended SAGConv and Topkpooling graph neural network (ST-GNN) to learn the features of each node in the tree. To enhance detection accuracy, we eliminate and merge some non-crucial nodes to highlight key nodes and execution orders. Finally, we evaluate our approach on 7, 962 real-world smart contracts running on Ethereum and compare it with state-of-the-art approaches on six types of vulnerabilities. Experimental results show that our approach achieves higher detection accuracy than others.

JMLR Journal 2017 Journal Article

COEVOLVE: A Joint Point Process Model for Information Diffusion and Network Evolution

  • Mehrdad Farajtabar
  • Yichen Wang
  • Manuel Gomez-Rodriguez
  • Shuang Li
  • Hongyuan Zha
  • Le Song

Information diffusion in online social networks is affected by the underlying network topology, but it also has the power to change it. Online users are constantly creating new links when exposed to new information sources, and in turn these links are alternating the way information spreads. However, these two highly intertwined stochastic processes, information diffusion and network evolution, have been predominantly studied separately, ignoring their co-evolutionary dynamics. We propose a temporal point process model, Coevolve, for such joint dynamics, allowing the intensity of one process to be modulated by that of the other. This model allows us to efficiently simulate interleaved diffusion and network events, and generate traces obeying common diffusion and network patterns observed in real-world networks such as Twitter. Furthermore, we also develop a convex optimization framework to learn the parameters of the model from historical diffusion and network evolution traces. We experimented with both synthetic data and data gathered from Twitter, and show that our model provides a good fit to the data as well as more accurate predictions than alternatives. [abs] [ pdf ][ bib ] &copy JMLR 2017. ( edit, beta )

NeurIPS Conference 2017 Conference Paper

Predicting User Activity Level In Point Processes With Mass Transport Equation

  • Yichen Wang
  • Xiaojing Ye
  • Hongyuan Zha
  • Le Song

Point processes are powerful tools to model user activities and have a plethora of applications in social sciences. Predicting user activities based on point processes is a central problem. However, existing works are mostly problem specific, use heuristics, or simplify the stochastic nature of point processes. In this paper, we propose a framework that provides an unbiased estimator of the probability mass function of point processes. In particular, we design a key reformulation of the prediction problem, and further derive a differential-difference equation to compute a conditional probability mass function. Our framework is applicable to general point processes and prediction tasks, and achieves superb predictive and efficiency performance in diverse real-world applications compared to state-of-arts.

NeurIPS Conference 2016 Conference Paper

Coevolutionary Latent Feature Processes for Continuous-Time User-Item Interactions

  • Yichen Wang
  • Nan Du
  • Rakshit Trivedi
  • Le Song

Matching users to the right items at the right time is a fundamental task in recommendation systems. As users interact with different items over time, users' and items' feature may evolve and co-evolve over time. Traditional models based on static latent features or discretizing time into epochs can become ineffective for capturing the fine-grained temporal dynamics in the user-item interactions. We propose a coevolutionary latent feature process model that accurately captures the coevolving nature of users' and items' feature. To learn parameters, we design an efficient convex optimization algorithm with a novel low rank space sharing constraints. Extensive experiments on diverse real-world datasets demonstrate significant improvements in user behavior prediction compared to state-of-the-arts.

NeurIPS Conference 2015 Conference Paper

COEVOLVE: A Joint Point Process Model for Information Diffusion and Network Co-evolution

  • Mehrdad Farajtabar
  • Yichen Wang
  • Manuel Gomez Rodriguez
  • Shuang Li
  • Hongyuan Zha
  • Le Song

Information diffusion in online social networks is affected by the underlying network topology, but it also has the power to change it. Online users are constantly creating new links when exposed to new information sources, and in turn these links are alternating the way information spreads. However, these two highly intertwined stochastic processes, information diffusion and network evolution, have been predominantly studied separately, ignoring their co-evolutionary dynamics. We propose a temporal point process model, COEVOLVE, for such joint dynamics, allowing the intensity of one process to be modulated by that of the other. This model allows us to efficiently simulate interleaved diffusion and network events, and generate traces obeying common diffusion and network patterns observed in real-world networks. Furthermore, we also develop a convex optimization framework to learn the parameters of the model from historical diffusion and network evolution traces. We experimented with both synthetic data and data gathered from Twitter, and show that our model provides a good fit to the data as well as more accurate predictions than alternatives.

IJCAI Conference 2015 Conference Paper

Detecting Emotions in Social Media: A Constrained Optimization Approach

  • Yichen Wang
  • Aditya Pal

Emotion detection can considerably enhance our understanding of users’ emotional states. Understanding users’ emotions especially in a real-time setting can be pivotal in improving user interactions and understanding their preferences. In this paper, we propose a constraint optimization framework to discover emotions from social media content of the users. Our framework employs several novel constraints such as emotion bindings, topic correlations, along with specialized features proposed by prior work and well-established emotion lexicons. We propose an efficient inference algorithm and report promising empirical results on three diverse datasets.

NeurIPS Conference 2015 Conference Paper

Time-Sensitive Recommendation From Recurrent User Activities

  • Nan Du
  • Yichen Wang
  • Niao He
  • Jimeng Sun
  • Le Song

By making personalized suggestions, a recommender system is playing a crucial role in improving the engagement of users in modern web-services. However, most recommendation algorithms do not explicitly take into account the temporal behavior and the recurrent activities of users. Two central but less explored questions are how to recommend the most desirable item \emph{at the right moment}, and how to predict \emph{the next returning time} of a user to a service. To address these questions, we propose a novel framework which connects self-exciting point processes and low-rank models to capture the recurrent temporal patterns in a large collection of user-item consumption pairs. We show that the parameters of the model can be estimated via a convex optimization, and furthermore, we develop an efficient algorithm that maintains $O(1 / \epsilon)$ convergence rate, scales up to problems with millions of user-item pairs and thousands of millions of temporal events. Compared to other state-of-the-arts in both synthetic and real datasets, our model achieves superb predictive performance in the two time-sensitive recommendation questions. Finally, we point out that our formulation can incorporate other extra context information of users, such as profile, textual and spatial features.

v2026.09.13