Arrow Research search

Author name cluster

Jan-Jakob Sonke

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
2 author rows

Possible papers

6

ICLR Conference 2025 Conference Paper

CaPo: Cooperative Plan Optimization for Efficient Embodied Multi-Agent Cooperation

  • Jie Liu 0043
  • Pan Zhou 0002
  • Yingjun Du
  • Ah-Hwee Tan
  • Cees G. M. Snoek
  • Jan-Jakob Sonke
  • Efstratios Gavves

In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods often execute actions extemporaneously and incoherently, without long-term strategic and cooperative planning, leading to redundant steps, failures, and even serious repercussions in complex tasks like search-and-rescue missions where discussion and cooperative plan are crucial. To solve this issue, we propose Cooperative Plan Optimization (CaPo) to enhance the cooperation efficiency of LLM-based embodied agents. Inspired by human cooperation schemes, CaPo improves cooperation efficiency with two phases: 1) meta plan generation, and 2) progress-adaptive meta plan and execution. In the first phase, all agents analyze the task, discuss, and cooperatively create a meta-plan that decomposes the task into subtasks with detailed steps, ensuring a long-term strategic and coherent plan for efficient coordination. In the second phase, agents execute tasks according to the meta-plan and dynamically adjust it based on their latest progress (e.g., discovering a target object) through multi-turn discussions. This progress-based adaptation eliminates redundant actions, improving the overall cooperation efficiency of agents. Experimental results on the ThreeDworld Multi-Agent Transport and Communicative Watch-And-Help tasks demonstrate CaPo's much higher task completion rate and efficiency compared with state-of-the-arts. The code is released at https://github.com/jliu4ai/CaPo.

ICML Conference 2025 Conference Paper

Probabilistic Interactive 3D Segmentation with Hierarchical Neural Processes

  • Jie Liu 0043
  • Pan Zhou 0002
  • Zehao Xiao
  • Jiayi Shen
  • Wenzhe Yin
  • Jan-Jakob Sonke
  • Efstratios Gavves

Interactive 3D segmentation has emerged as a promising solution for generating accurate object masks in complex 3D scenes by incorporating user-provided clicks. However, two critical challenges remain underexplored: (1) effectively generalizing from sparse user clicks to produce accurate segmentations and (2) quantifying predictive uncertainty to help users identify unreliable regions. In this work, we propose NPISeg3D, a novel probabilistic framework that builds upon Neural Processes (NPs) to address these challenges. Specifically, NPISeg3D introduces a hierarchical latent variable structure with scene-specific and object-specific latent variables to enhance few-shot generalization by capturing both global context and object-specific characteristics. Additionally, we design a probabilistic prototype modulator that adaptively modulates click prototypes with object-specific latent variables, improving the model’s ability to capture object-aware context and quantify predictive uncertainty. Experiments on four 3D point cloud datasets demonstrate that NPISeg3D achieves superior segmentation performance with fewer clicks while providing reliable uncertainty estimations.

UAI Conference 2024 Conference Paper

Domain Adaptation with Cauchy-Schwarz Divergence

  • Wenzhe Yin
  • Shujian Yu
  • Yicong Lin
  • Jie Liu 0043
  • Jan-Jakob Sonke
  • Efstratios Gavves

Domain adaptation aims to use training data from one or multiple source domains to learn a hypothesis that can be generalized to a different, but related, target domain. As such, having a reliable measure for evaluating the discrepancy of both marginal and conditional distributions is crucial. We introduce Cauchy-Schwarz (CS) divergence to the problem of unsupervised domain adaptation (UDA). The CS divergence offers a theoretically tighter generalization error bound than the popular Kullback-Leibler divergence. This holds for the general case of supervised learning, including multi-class classification and regression. Furthermore, we illustrate that the CS divergence enables a simple estimator on the discrepancy of both marginal and conditional distributions between source and target domains in the representation space, without requiring any distributional assumptions. We provide multiple examples to illustrate how the CS divergence can be conveniently used in both distance metric- or adversarial training-based UDA frameworks, resulting in compelling performance. The code of our paper is available at \url{https: //github. com/ywzcode/CS-adv}.

NeurIPS Conference 2024 Conference Paper

Space-Time Continuous PDE Forecasting using Equivariant Neural Fields

  • David M. Knigge
  • David R. Wessels
  • Riccardo Valperga
  • Samuele Papa
  • Jan-Jakob Sonke
  • Efstratios Gavves
  • Erik J. Bekkers

Recently, Conditional Neural Fields (NeFs) have emerged as a powerful modelling paradigm for PDEs, by learning solutions as flows in the latent space of the Conditional NeF. Although benefiting from favourable properties of NeFs such as grid-agnosticity and space-time-continuous dynamics modelling, this approach limits the ability to impose known constraints of the PDE on the solutions -- such as symmetries or boundary conditions -- in favour of modelling flexibility. Instead, we propose a space-time continuous NeF-based solving framework that - by preserving geometric information in the latent space of the Conditional NeF - preserves known symmetries of the PDE. We show that modelling solutions as flows of pointclouds over the group of interest $G$ improves generalization and data-efficiency. Furthermore, we validate that our framework readily generalizes to unseen spatial and temporal locations, as well as geometric transformations of the initial conditions - where other NeF-based PDE forecasting methods fail -, and improve over baselines in a number of challenging geometries.

ICLR Conference 2023 Conference Paper

Modelling Long Range Dependencies in $N$D: From Task-Specific to a General Purpose CNN

  • David M. Knigge
  • David W. Romero
  • Albert Gu
  • Efstratios Gavves
  • Erik J. Bekkers
  • Jakub M. Tomczak
  • Mark Hoogendoorn
  • Jan-Jakob Sonke

Performant Convolutional Neural Network (CNN) architectures must be tailored to specific tasks in order to consider the length, resolution, and dimensionality of the input data. In this work, we tackle the need for problem-specific CNN architectures. We present the Continuous Convolutional Neural Network (CCNN): a single CNN able to process data of arbitrary resolution, dimensionality and length without any structural changes. Its key component are its continuous convolutional kernels which model long-range dependencies at every layer, and thus remove the need of current CNN architectures for task-dependent downsampling and depths. We showcase the generality of our method by using the same architecture for tasks on sequential ($1{\rm D}$), visual ($2{\rm D}$) and point-cloud ($3{\rm D}$) data. Our CCNN matches and often outperforms the current state-of-the-art across all tasks considered.

YNIMG Journal 2022 Journal Article

A unified model for reconstruction and R 2 * mapping of accelerated 7T data using the quantitative recurrent inference machine

  • Chaoping Zhang
  • Dimitrios Karkalousos
  • Pierre-Louis Bazin
  • Bram F. Coolen
  • Hugo Vrenken
  • Jan-Jakob Sonke
  • Birte U. Forstmann
  • Dirk H.J. Poot

Quantitative MRI (qMRI) acquired at the ultra-high field of 7 Tesla has been used in visualizing and analyzing subcortical structures. qMRI relies on the acquisition of multiple images with different scan settings, leading to extended scanning times. Data redundancy and prior information from the relaxometry model can be exploited by deep learning to accelerate the imaging process. We propose the quantitative Recurrent Inference Machine (qRIM), with a unified forward model for joint reconstruction and R 2 * -mapping from sparse data, embedded in a Recurrent Inference Machine (RIM), an iterative inverse problem-solving network. To study the dependency of the proposed extension of the unified forward model to network architecture, we implemented and compared a quantitative End-to-End Variational Network (qE2EVN). Experiments were performed with high-resolution multi-echo gradient echo data of the brain at 7T of a cohort study covering the entire adult life span. The error in reconstructed R 2 * from undersampled data relative to reference data significantly decreased for the unified model compared to sequential image reconstruction and parameter fitting using the RIM. With increasing acceleration factor, an increasing reduction in the reconstruction error was observed, pointing to a larger benefit for sparser data. Qualitatively, this was following an observed reduction of image blurriness in R 2 * -maps. In contrast, when using the U-Net as network architecture, a negative bias in R 2 * in selected regions of interest was observed. Compressed Sensing rendered accurate, but less precise estimates of R 2 *. The qE2EVN showed slightly inferior reconstruction quality compared to the qRIM but better quality than the U-Net and Compressed Sensing. Subcortical maturation over age measured by a linearly increasing interquartile range of R 2 * in the striatum was preserved up to an acceleration factor of 9. With the integrated prior of the unified forward model, the proposed qRIM can exploit the redundancy among repeated measurements and shared information between tasks, facilitating relaxometry in accelerated MRI.

v2026.09.13