Arrow Research search

Author name cluster

Howard Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

NeurIPS Conference 2023 Conference Paper

Fed-GraB: Federated Long-tailed Learning with Self-Adjusting Gradient Balancer

  • Zikai Xiao
  • Zihan Chen
  • Songshang Liu
  • Hualiang Wang
  • Yang Feng
  • Jin Hao
  • Joey Tianyi Zhou
  • Jian Wu

Data privacy and long-tailed distribution are the norms rather than the exception in many real-world tasks. This paper investigates a federated long-tailed learning (Fed-LT) task in which each client holds a locally heterogeneous dataset; if the datasets can be globally aggregated, they jointly exhibit a long-tailed distribution. Under such a setting, existing federated optimization and/or centralized long-tailed learning methods hardly apply due to challenges in (a) characterizing the global long-tailed distribution under privacy constraints and (b) adjusting the local learning strategy to cope with the head-tail imbalance. In response, we propose a method termed $\texttt{Fed-GraB}$, comprised of a Self-adjusting Gradient Balancer (SGB) module that re-weights clients' gradients in a closed-loop manner, based on the feedback of global long-tailed distribution evaluated by a Direct Prior Analyzer (DPA) module. Using $\texttt{Fed-GraB}$, clients can effectively alleviate the distribution drift caused by data heterogeneity during the model training process and obtain a global model with better performance on the minority classes while maintaining the performance of the majority classes. Extensive experiments demonstrate that $\texttt{Fed-GraB}$ achieves state-of-the-art performance on representative datasets such as CIFAR-10-LT, CIFAR-100-LT, ImageNet-LT, and iNaturalist.

IROS Conference 2023 Conference Paper

Object-Oriented Option Framework for Robotics Manipulation in Clutter

  • Jing-Cheng Pang
  • Si-Hang Yang
  • Xiong-Hui Chen
  • Xinyu Yang
  • Yang Yu 0001
  • Mas Ma
  • Ziqi Guo
  • Howard Yang

Domestic service robots are becoming increasingly popular due to their ability to help people with household tasks. These robots often encounter the challenge of manipulating objects in cluttered environments (MoC), which is difficult due to the complexity of effective planning and control. Previous solutions involved designing specific action primitives and planning paradigms. However, the pre-coded action primitives can limit the agility and task-solving scope of robots. In this paper, we propose a general approach for MoC called the Object-Oriented Option Framework (O3F), which uses the option framework (OF) to learn planning and control. The standard OF discovers options from scratch based on reinforcement learning, which can lead to collapsed options and hurt learning. To address this limitation, O3F introduces the concept of an object-oriented option space for OF, which focuses specifically on object movement and overcomes the challenges associated with collapsed options. Based on this, we train an object-oriented option planner to determine the option to execute and a universal object-oriented option executor to complete the option. Simulation experiments on the Ginger XR1 robot and robot arm show that O3F is generally applicable to various types of robot and manipulation tasks. Furthermore, O3F achieves success rates of 72. 4% and 90% in grasping and object collecting tasks, respectively, significantly outperforming baseline methods.

NeurIPS Conference 2023 Conference Paper

Spectral Co-Distillation for Personalized Federated Learning

  • Zihan Chen
  • Howard Yang
  • Tony Quek
  • Kai Fong Ernest Chong

Personalized federated learning (PFL) has been widely investigated to address the challenge of data heterogeneity, especially when a single generic model is inadequate in satisfying the diverse performance requirements of local clients simultaneously. Existing PFL methods are inherently based on the idea that the relations between the generic global and personalized local models are captured by the similarity of model weights. Such a similarity is primarily based on either partitioning the model architecture into generic versus personalized components or modeling client relationships via model weights. To better capture similar (yet distinct) generic versus personalized model representations, we propose $\textit{spectral distillation}$, a novel distillation method based on model spectrum information. Building upon spectral distillation, we also introduce a co-distillation framework that establishes a two-way bridge between generic and personalized model training. Moreover, to utilize the local idle time in conventional PFL, we propose a wait-free local training protocol. Through extensive experiments on multiple datasets over diverse heterogeneous data settings, we demonstrate the outperformance and efficacy of our proposed spectral co-distillation method, as well as our wait-free training protocol.

NeurIPS Conference 1999 Conference Paper

Data Visualization and Feature Selection: New Algorithms for Nongaussian Data

  • Howard Yang
  • John Moody

Data visualization and feature selection methods are proposed based on the )oint mutual information and ICA. The visualization methods can find many good 2-D projections for high dimensional data interpretation, which cannot be easily found by the other ex(cid: 173) isting methods. The new variable selection method is found to be better in eliminating redundancy in the inputs than other methods based on simple mutual information. The efficacy of the methods is illustrated on a radar signal analysis problem to find 2-D viewing coordinates for data visualization and to select inputs for a neural network classifier. Keywords: feature selection, joint mutual information, ICA, vi(cid: 173) sualization, classification.

NeurIPS Conference 1999 Conference Paper

Search for Information Bearing Components in Speech

  • Howard Yang
  • Hynek Hermansky

In this paper, we use mutual information to characterize the dis(cid: 173) tributions of phonetic and speaker/channel information in a time(cid: 173) frequency space. The mutual information (MI) between the pho(cid: 173) netic label and one feature, and the joint mutual information (JMI) between the phonetic label and two or three features are estimated. The Miller's bias formulas for entropy and mutual information es(cid: 173) timates are extended to include higher order terms. The MI and the JMI for speaker/channel recognition are also estimated. The results are complementary to those for phonetic classification. Our results show how the phonetic information is locally spread and how the speaker/channel information is globally spread in time and frequency.

NeurIPS Conference 1997 Conference Paper

Multiplicative Updating Rule for Blind Separation Derived from the Method of Scoring

  • Howard Yang

For blind source separation, when the Fisher information matrix is used as the Riemannian metric tensor for the parameter space, the steepest descent algorithm to maximize the likelihood function in this Riemannian parameter space becomes the serial updating rule with equivariant property. This algorithm can be further simplified by using the asymptotic form of the Fisher information matrix around the equilibrium.

NeurIPS Conference 1997 Conference Paper

The Efficiency and the Robustness of Natural Gradient Descent Learning Rule

  • Howard Yang
  • Shun-ichi Amari

The inverse of the Fisher information matrix is used in the natu(cid: 173) ral gradient descent algorithm to train single-layer and multi-layer perceptrons. We have discovered a new scheme to represent the Fisher information matrix of a stochastic multi-layer perceptron. Based on this scheme, we have designed an algorithm to compute the natural gradient. When the input dimension n is much larger than the number of hidden neurons, the complexity of this algo(cid: 173) rithm is of order O(n). It is confirmed by simulations that the natural gradient descent learning rule is not only efficient but also robust.

NeurIPS Conference 1995 Conference Paper

A New Learning Algorithm for Blind Signal Separation

  • Shun-ichi Amari
  • Andrzej Cichocki
  • Howard Yang

A new on-line learning algorithm which minimizes a statistical de(cid: 173) pendency among outputs is derived for blind separation of mixed signals. The dependency is measured by the average mutual in(cid: 173) formation (MI) of the outputs. The source signals and the mixing matrix are unknown except for the number of the sources. The Gram-Charlier expansion instead of the Edgeworth expansion is used in evaluating the MI. The natural gradient approach is used to minimize the MI. A novel activation function is proposed for the on-line learning algorithm which has an equivariant property and is easily implemented on a neural network like model. The validity of the new learning algorithm are verified by computer simulations.

NeurIPS Conference 1995 Conference Paper

Statistical Theory of Overtraining - Is Cross-Validation Asymptotically Effective?

  • Shun-ichi Amari
  • Noboru Murata
  • Klaus-Robert Müller
  • Michael Finke
  • Howard Yang

A statistical theory for overtraining is proposed. The analysis treats realizable stochastic neural networks, trained with Kullback(cid: 173) Leibler loss in the asymptotic case. It is shown that the asymptotic gain in the generalization error is small if we perform early stop(cid: 173) ping, even if we have access to the optimal stopping time. Consider(cid: 173) ing cross-validation stopping we answer the question: In what ratio the examples should be divided into training and testing sets in or(cid: 173) der to obtain the optimum performance. In the non-asymptotic region cross-validated early stopping always decreases the general(cid: 173) ization error. Our large scale simulations done on a CM5 are in nice agreement with our analytical findings.

v2026.09.13