Arrow Research search

Author name cluster

Yoshitaka Ushiku

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

13 papers
2 author rows

Possible papers

13

ICLR Conference 2025 Conference Paper

Rethinking the role of frames for SE(3)-invariant crystal structure modeling

  • Yusei Ito
  • Tatsunori Taniai
  • Ryo Igarashi 0002
  • Yoshitaka Ushiku
  • Kanta Ono

Crystal structure modeling with graph neural networks is essential for various applications in materials informatics, and capturing SE(3)-invariant geometric features is a fundamental requirement for these networks. A straightforward approach is to model with orientation-standardized structures through structure-aligned coordinate systems, or “frames.” However, unlike molecules, determining frames for crystal structures is challenging due to their infinite and highly symmetric nature. In particular, existing methods rely on a statically fixed frame for each structure, determined solely by its structural information, regardless of the task under consideration. Here, we rethink the role of frames, *questioning whether such simplistic alignment with the structure is sufficient*, and propose the concept of *dynamic frames*. While accommodating the infinite and symmetric nature of crystals, these frames provide each atom with a dynamic view of its local environment, focusing on actively interacting atoms. We demonstrate this concept by utilizing the attention mechanism in a recent transformer-based crystal encoder, resulting in a new architecture called **CrystalFramer**. Extensive experiments show that CrystalFramer outperforms conventional frames and existing crystal encoders in various crystal property prediction tasks.

ICRA Conference 2025 Conference Paper

SCU-Hand: Soft Conical Universal Robotic Hand for Scooping Granular Media from Containers of Various Sizes

  • Tomoya Takahashi
  • Cristian C. Beltran-Hernandez
  • Yuki Kuroda
  • Kazutoshi Tanaka
  • Masashi Hamaya
  • Yoshitaka Ushiku

Automating small-scale experiments in materials science presents challenges due to the heterogeneous nature of experimental setups. This study introduces the SCU-Hand (Soft Conical Universal Robot Hand), a novel end-effector designed to automate the task of scooping powdered samples from various container sizes using a robotic arm. The SCU-Hand employs a flexible, conical structure that adapts to different container geometries through deformation, maintaining consistent contact without complex force sensing or machine learning-based control methods. Its reconfigurable mechanism allows for size adjustment, enabling efficient scooping from diverse container types. By combining soft robotics principles with a sheet-morphing design, our end-effector achieves high flexibility while retaining the necessary stiffness for effective powder manipulation. We detail the design principles, fabrication process, and experimental validation of the SCU-Hand. Experimental validation showed that the scooping capacity is about 20% higher than that of a commercial tool, with a scooping performance of more than 95% for containers of sizes between 67 mm to 110 mm. This research contributes to laboratory automation by offering a cost-effective, easily implementable solution for automating tasks such as materials synthesis and characterization processes.

ICLR Conference 2024 Conference Paper

Crystalformer: Infinitely Connected Attention for Periodic Structure Encoding

  • Tatsunori Taniai
  • Ryo Igarashi 0002
  • Yuta Suzuki
  • Naoya Chiba
  • Kotaro Saito
  • Yoshitaka Ushiku
  • Kanta Ono

Predicting physical properties of materials from their crystal structures is a fundamental problem in materials science. In peripheral areas such as the prediction of molecular properties, fully connected attention networks have been shown to be successful. However, unlike these finite atom arrangements, crystal structures are infinitely repeating, periodic arrangements of atoms, whose fully connected attention results in *infinitely connected attention*. In this work, we show that this infinitely connected attention can lead to a computationally tractable formulation, interpreted as *neural potential summation*, that performs infinite interatomic potential summations in a deeply learned feature space. We then propose a simple yet effective Transformer-based encoder architecture for crystal structures called *Crystalformer*. Compared to an existing Transformer-based model, the proposed model requires only 29.4% of the number of parameters, with minimal modifications to the original Transformer architecture. Despite the architectural simplicity, the proposed method outperforms state-of-the-art methods for various property regression tasks on the Materials Project and JARVIS-DFT datasets.

ICRA Conference 2024 Conference Paper

Vision-Language Interpreter for Robot Task Planning

  • Keisuke Shirai
  • Cristian C. Beltran-Hernandez
  • Masashi Hamaya
  • Atsushi Hashimoto 0001
  • Shohei Tanaka
  • Kento Kawaharazuka
  • Kazutoshi Tanaka
  • Yoshitaka Ushiku

Large language models (LLMs) are accelerating the development of language-guided robot planners. Meanwhile, symbolic planners offer the advantage of interpretability. This paper proposes a new task that bridges these two trends, namely, multimodal planning problem specification. The aim is to generate a problem description (PD), a machine-readable file used by the planners to find a plan. By generating PDs from language instruction and scene observation, we can drive symbolic planners in a language-guided framework. We propose a Vision-Language Interpreter (ViLaIn), a new framework that generates PDs using state-of-the-art LLM and vision-language models. ViLaIn can refine generated PDs via error message feedback from the symbolic planner. Our aim is to answer the question: How accurately can ViLaIn and the symbolic planner generate valid robot plans? To evaluate ViLaIn, we introduce a novel dataset called the problem description generation (ProDG) dataset. The framework is evaluated with four new evaluation metrics. Experimental results show that ViLaIn can generate syntactically correct problems with more than 99% accuracy and valid plans with more than 58% accuracy. Our code and dataset are available at https://github.com/omron-sinicx/ViLaIn.

IROS Conference 2023 Conference Paper

Robotic Powder Grinding with Audio-Visual Feedback for Laboratory Automation in Materials Science

  • Yusaku Nakajima
  • Masashi Hamaya
  • Kazutoshi Tanaka
  • Takafumi Hawai
  • Felix von Drigalski
  • Yasuo Takeichi
  • Yoshitaka Ushiku
  • Kanta Ono

This study focuses on the powder grinding process, which is a necessary step for material synthesis in materials science experiments. In material science, powder grinding is a time-consuming process that is typically executed by hand, as commercial grinding machines are unsuitable for samples of small size. Robotic powder grinding would solve this problem, but it is a challenging task for robots, as it requires observing the powder state and generating appropriate motions. Our previous study proposed a robotic powder grinding system using visual feedback. Although visual feedback is helpful for observing the powder distribution, the particle size during the grinding process remains invisible, leading to suboptimal robot actions. In some cases, the robot chose to gather the powder even though continuing to grind instead would have produced finer powder. In this paper, we present a multi-modal robotic grinding system that utilizes both audio and visual feedback. It makes use of the grinding sound which carries information about the grinding progress, as the particle size strongly affects the audio intensity. The audio feedback enables the robot to grind until the powder is sufficiently fine. In our experiments, the robot ground 80. 5% of the powder to a particle size smaller than $250\ \mu\mathrm{m}$ with audio and visual feedback and 68% without audio feedback, indicating that multi-modal feedback is an effective tool to produce finer powder. We conclude that the addition of audio feedback provides crucial information to the robot, allowing it to better understand the progress of the grinding process and make more optimal decisions. This robot system can be used to prepare samples in material science experiments and analyze the grinding process.

IROS Conference 2022 Conference Paper

Robotic Powder Grinding with a Soft Jig for Laboratory Automation in Material Science

  • Yusaku Nakajima
  • Masashi Hamaya
  • Yuta Suzuki
  • Takafumi Hawai
  • Felix von Drigalski
  • Kazutoshi Tanaka
  • Yoshitaka Ushiku
  • Kanta Ono

Grinding materials into a fine powder is a time-consuming task in material science that is generally performed by hand, as current automated grinding machines might not be suitable for preparing small-sized samples. This study presents a robotic powder grinding system for laboratory automation in material science applications that observe the powder's state to improve the grinding outcome. We developed a soft jig consisting of off-the-shelf gel materials and 3D-printed parts, which can be used with any robot arm to perform powder grinding. The jig's physical softness allows for safe grinding without force sensing. In addition, we developed a visual feedback system that observes the powder distribution and decides where to grind and when to gather. The results showed that our system could grind 79 percent of the powder to a particle size smaller than 200 μm by using the soft jig and visual feedback. This ratio was 57% when using only the soft jig without feedback. Our system can be used immediately in laboratories to alleviate the workload of researchers.

AAAI Conference 2019 Conference Paper

Estimating the Causal Effect from Partially Observed Time Series

  • Akane Iseki
  • Yusuke Mukuta
  • Yoshitaka Ushiku
  • Tatsuya Harada

Many real-world systems involve interacting time series. The ability to detect causal dependencies between system components from observed time series of their outputs is essential for understanding system behavior. The quantification of causal influences between time series is based on the definition of some causality measure. Partial Canonical Correlation Analysis (Partial CCA) and its extensions are examples of methods used for robustly estimating the causal relationships between two multidimensional time series even when the time series are short. These methods assume that the input data are complete and have no missing values. However, real-world data often contain missing values. It is therefore crucial to estimate the causality measure robustly even when the input time series is incomplete. Treating this problem as a semi-supervised learning problem, we propose a novel semi-supervised extension of probabilistic Partial CCA called semi-Bayesian Partial CCA. Our method exploits the information in samples with missing values to prevent the overfitting of parameter estimation even when there are few complete samples. Experiments based on synthesized and real data demonstrate the ability of the proposed method to estimate causal relationships more correctly than existing methods when the data contain missing values, the dimensionality is large, and the number of samples is small.

ICRA Conference 2019 Conference Paper

Pose Graph optimization for Unsupervised Monocular Visual Odometry

  • Yang Li 0143
  • Yoshitaka Ushiku
  • Tatsuya Harada

Unsupervised Learning based monocular visual odometry (VO) has lately drawn significant attention for its potential in label-free leaning ability and robustness to camera parameters and environmental variations. However, partially due to the lack of drift correction technique, these methods are still by far less accurate than geometric approaches for large-scale odometry estimation. In this paper, we propose to leverage graph optimization and loop closure detection to overcome limitations of unsupervised learning based monocular visual odometry. To this end, we propose a hybrid VO system which combines an unsupervised monocular VO called NeuralBundler with a pose graph optimization back-end. NeuralBundler is a neural network architecture that uses temporal and spatial photometric loss as main supervision and generates a windowed pose graph consists of multi-view 6DoF constraints. We propose a novel pose cycle consistency loss to relieve the tensions in the windowed pose graph, leading to improved performance and robustness. In the back-end, a global pose graph is built from local and loop 6DoF constraints estimated by NeuralBundler, and is optimized over SE(3). Empirical evaluation on the KITTI odometry dataset demonstrates that 1) NeuralBundler achieves state-of-the-art performance on unsupervised monocular VO estimation, and 2) our whole approach can achieve efficient loop closing and show favorable overall translational accuracy compared to established monocular SLAM systems.

AAAI Conference 2018 Conference Paper

Alternating Circulant Random Features for Semigroup Kernels

  • Yusuke Mukuta
  • Yoshitaka Ushiku
  • Tatsuya Harada

The random features method is an efficient method to approximate the kernel function. In this paper, we propose novel random features called “alternating circulant random features, ” which consist of a random mixture of independent random structured matrices. Existing fast random features exploit random sign flipping to reduce the correlation between features. Sign flipping works well on random Fourier features for real-valued shift-invariant kernels because the corresponding weight distribution is symmetric. However, this method cannot be applied to random Laplace features directly because the distribution is not symmetric. The method proposed herein yields alternating circulant random features, with the correlation between features being reduced through the random sampling of weights from multiple independent random structured matrices instead of via random sign flipping. The proposed method facilitates rapid calculation by employing structured matrices. In addition, the weight distribution is preserved because sign flipping is not implemented. The performance of the proposed alternating circulant random features method is theoretically and empirically evaluated.

AAAI Conference 2018 Conference Paper

Hierarchical Video Generation From Orthogonal Information: Optical Flow and Texture

  • Katsunori Ohnishi
  • Shohei Yamamoto
  • Yoshitaka Ushiku
  • Tatsuya Harada

Learning to represent and generate videos from unlabeled data is a very challenging problem. To generate realistic videos, it is important not only to ensure that the appearance of each frame is real, but also to ensure the plausibility of a video motion and consistency of a video appearance in the time direction. The process of video generation should be divided according to these intrinsic difficulties. In this study, we focus on the motion and appearance information as two important orthogonal components of a video, and propose Flow-and-Texture-Generative Adversarial Networks (FTGAN) consisting of FlowGAN and TextureGAN. In order to avoid a huge annotation cost, we have to explore a way to learn from unlabeled data. Thus, we employ optical flow as motion information to generate videos. FlowGAN generates optical flow, which contains only the edge and motion of the videos to be begerated. On the other hand, Texture- GAN specializes in giving a texture to optical flow generated by FlowGAN. This hierarchical approach brings more realistic videos with plausible motion and appearance consistency. Our experiments show that our model generates more plausible motion videos and also achieves significantly improved performance for unsupervised action classification in comparison to previous GAN works. In addition, because our model generates videos from two independent information, our model can generate new combinations of motion and attribute that are not seen in training data, such as a video in which a person is doing sit-up in a baseball ground.

ICML Conference 2017 Conference Paper

Asymmetric Tri-training for Unsupervised Domain Adaptation

  • Kuniaki Saito
  • Yoshitaka Ushiku
  • Tatsuya Harada

It is important to apply models trained on a large number of labeled samples to different domains because collecting many labeled samples in various domains is expensive. To learn discriminative representations for the target domain, we assume that artificially labeling the target samples can result in a good representation. Tri-training leverages three classifiers equally to provide pseudo-labels to unlabeled samples; however, the method does not assume labeling samples generated from a different domain. In this paper, we propose the use of an asymmetric tri-training method for unsupervised domain adaptation, where we assign pseudo-labels to unlabeled samples and train the neural networks as if they are true labels. In our work, we use three networks asymmetrically, and by asymmetric, we mean that two networks are used to label unlabeled target samples, and one network is trained by the pseudo-labeled samples to obtain target-discriminative representations. Our proposed method was shown to achieve a state-of-the-art performance on the benchmark digit recognition datasets for domain adaptation.

IROS Conference 2017 Conference Paper

MFNet: Towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes

  • Qishen Ha
  • Kohei Watanabe
  • Takumi Karasawa
  • Yoshitaka Ushiku
  • Tatsuya Harada

This work addresses the semantic segmentation of images of street scenes for autonomous vehicles based on a new RGB-Thermal dataset, which is also introduced in this paper. An increasing interest in self-driving vehicles has brought the adaptation of semantic segmentation to self-driving systems. However, recent research relating to semantic segmentation is mainly based on RGB images acquired during times of poor visibility at night and under adverse weather conditions. Furthermore, most of these methods only focused on improving performance while ignoring time consumption. The aforementioned problems prompted us to propose a new convolutional neural network architecture for multi-spectral image segmentation that enables the segmentation accuracy to be retained during real-time operation. We benchmarked our method by creating an RGB-Thermal dataset in which thermal and RGB images are combined. We showed that the segmentation accuracy was significantly increased by adding thermal infrared information.

ICRA Conference 2014 Conference Paper

Hard negative classes for multiple object detection

  • Asako Kanezaki
  • Sho Inaba
  • Yoshitaka Ushiku
  • Yuya Yamashita
  • Hiroshi Muraoka
  • Yasuo Kuniyoshi
  • Tatsuya Harada

We propose an efficient method to train multiple object detectors simultaneously using a large scale image dataset. The one-vs-all approach that optimizes the boundary between positive samples from a target class and negative samples from the others has been the most standard approach for object detection. However, because this approach trains each object detector independently, the scores are not balanced between object classes. The proposed method combines ideas derived from both detection and classification in order to balance the scores across all object classes. We optimized the boundary between target classes and their “hard negative” samples, just as in detection, while simultaneously balancing the detector scores across object classes, as done in multi-class classification. We evaluated the performances on multi-class object detection using a subset of the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) 2011 dataset and showed our method outperformed a de facto standard method.

v2026.09.13