Arrow Research search

Author name cluster

Leszek Rutkowski

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
1 author row

Possible papers

5

AAAI Conference 2026 Conference Paper

Bridging the Tokenizer Gap: Semantics and Distribution-aware Knowledge Transfer for Unbiased Cross-Tokenizer Distillation

  • Huazheng Wang
  • Yongcheng Jing
  • Haifeng Sun
  • Jingyu Wang
  • Jianxin Liao
  • Leszek Rutkowski
  • Dacheng Tao

Cross-tokenizer knowledge distillation, where the teacher and student employ different tokenizers, is becoming increasingly prevalent, yet it poses underexplored challenges: existing methods fail to capture the rich knowledge encoded in teacher logits, as evidenced by the neglect of semantic information, inaccurate and biased logit alignment, and discarding distributional structure—ultimately leading to unfavorable distillation. To address these issues, we propose SeDi, a semantics and distribution-aware knowledge transfer framework tailored for cross-tokenizer distillation. To preserve factual knowledge, SeDi employs bipartite graph-based alignment at the tokenization level and a sliding window re-encoding strategy at the vocabulary level, enabling unbiased transfer of the teacher’s next-token predictions into the student’s vocabulary space. To further retain distributional information, we align the student’s entropy with that of the teacher by incorporating the student’s own logits during training, which helps to mitigate the exposure bias problem. Experiments on ten datasets across three task domains and five different teacher-student model pairs with varying vocabulary sizes demonstrate that SeDi delivers substantial improvements, with gains of up to 19.8%.

AAAI Conference 2026 Conference Paper

SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMs

  • Shuhan Xu
  • Siyuan Liang
  • Hongling Zheng
  • Aishan Liu
  • Xinbiao Wang
  • Yong Luo
  • Fu Lin
  • Leszek Rutkowski

Visual language models (VLMs) have made significant progress in image captioning tasks, yet recent studies have found they are vulnerable to backdoor attacks. Attackers can inject undetectable perturbations into the data during inference, triggering abnormal behavior and generating malicious captions. These attacks are particularly challenging to detect and defend against due to the stealthiness and cross-modal propagation of the trigger signals. In this paper, we identify two key vulnerabilities by analyzing existing attack patterns: (1) the model exhibits abnormal attention concentration on certain regions of the input image, and (2) backdoor attacks often induce semantic drift and sentence incoherence. Based on these insights, we propose Semantic Reward Defense (SRD), a reinforcement learning framework that mitigates backdoor behavior without requiring any prior knowledge of trigger patterns. SRD learns to apply discrete perturbations to sensitive contextual regions of image inputs via a deep Q-network policy, aiming to confuse attention and disrupt the activation of malicious paths. To guide policy optimization, we design a reward signal named semantic fidelity score, which jointly assesses the semantic consistency and linguistic fluency of the generated captions, encouraging the agent to achieve a robust yet faithful output. SRD offers a trigger-agnostic, policy-interpretable defense paradigm that effectively mitigates local (TrojVLM) and global (Shadowcast) backdoor attacks, reducing ASR to 3.6% and 5.6% respectively, with less than 15% average CIDEr drop on the clean inputs.

AAAI Conference 2026 Conference Paper

Time-Frequency Token Advantage Clipping for Training Efficient Large Reasoning Model

  • Rong Bao
  • Bo Wang
  • Xiao Wang
  • Hongyu Li
  • Rui Zheng
  • Leszek Rutkowski
  • Qi Zhang
  • Liang Ding

Long Chain-of-Thought (CoT) reasoning enhances large reasoning models' performance but suffers from severe inefficiencies, as models often overthink simple problems or underthink complex ones. Current sequence-level optimizations, like length penalties, are too coarse-grained to distinguish core logic from verbose language, precluding the necessary token-level control for efficient reasoning CoT. To overcome these limitations, we introduce Time-Frequency token Advantage Clipping (TFAC), a novel training framework designed to build efficient large reasoning models via token-level interventions. Specifically, TFAC functions along two dimensions: 1) The Frequency Dimension: It discourages inefficient loops and encourages deeper exploration by dynamically reducing the advantage scores of high-entropy tokens that are repeatedly generated within a single reasoning path. 2) The Time Dimension: It reduces excessive overthinking of the system by establishing a historical baseline for the occurrence count of each critical token in previously successful trajectories, and clipping the advantages of tokens that exceed this baseline during training. Crucially, to preserve the model's exploratory capabilities on novel problems, this suppression mechanism is automatically disabled when no historical record of success is available. Experiments conducted on the Deepseek-Distill-32B and Qwen3-8B models show that TFAC outperforms leading baseline methods, improving performance by 2.3 and 3.1 percentage points, respectively, while simultaneously reducing inference costs by 35% and 28% in scenarios where correct answers are generated. These results validate the significant efficacy of TFAC in training large reasoning models that are both powerful and highly efficient.

EAAI Journal 2025 Journal Article

A fuzzy three-criteria optimization-based currency trading system with adaptive criteria shapes and money management

  • Krzysztof Kaczmarek
  • Pavel Sevastjanov
  • Ludmila Dymova
  • Adam Kulawik
  • Leszek Rutkowski

The paper is focused on the development of algorithmic trading modeling — the modern branch of financial engineering. The concepts of negative overfitting risk and irregular shape of fuzzy criteria in currency trading modeling have been introduced. The first concept is based on observations, according to which deep drawbacks of the profit curve at the optimization stage can lead to negative overfitting (loss of profit during the testing period), as well as high local peaks on this profit curve. Thus, we can say that any significant irregularity of the profit curve can lead to losses due to the emergence of the effect of overfitting. To implement this concept, two relevant indicators of technical analysis and, based on them, risk criteria were developed. The second concept is based on the informal but logically justified assumption that irregular optimized forms of used membership functions may better reflect current market conditions. To implement these concepts, two fuzzy three-criteria optimization-based currency trading models comprised the profitability, reliability and risk local criteria with adaptive criteria shapes and money management were developed. These models were confirmed by the results obtained by modeling intraday trading over the past five years and four months on a 1-hour timeframe and the six most traded currency pairs. A comparison of these results with the results obtained by other authors confirmed their significant superiority and thus confirmed the correctness of the concepts introduced.

EAAI Journal 2023 Journal Article

A novel approach to intelligent monitoring of gas composition and light mode of greenhouse crop growing zone on the basis of fuzzy modelling and human-in-the-loop techniques

  • Ivan Laktionov
  • Leszek Rutkowski
  • Oleksandr Vovna
  • Aleksander Byrski
  • Maryna Kabanets

Gas composition and light mode of industrial greenhouses are some of the most determining factors in the process of growing vegetable crops in greenhouse conditions. The intelligentisation of information technologies for monitoring and control based on artificial intelligence methods can increase the efficiency of agrotechnical procedures for greenhouse cultivation. One of the efficient approaches in today's world practice is the development and implementation of trustworthy hybrid decision-support systems based on the techniques of Fuzzy logic and Human-in-the-Loop. The research object is non-stationary processes of complex intelligent transformation of measurement data on the concentration of carbon dioxide and effective energy illumination in the growing zone of industrial greenhouses. The scientific novelty and practical value of the obtained research results consist in creating a computer model for aggregation and intelligent processing of agricultural monitoring data for greenhouses. The developed computer model is fully transformed into peripheral-level software of Internet of Things systems for agricultural purposes. This makes it possible to implement the for-computing architecture of monitoring systems in greenhouses. The obtained research results make it possible to optimise the resources used in growing crops in greenhouse conditions through the implementation of hardware and software intelligent monitoring tools that are adaptive to the types and periods of crop vegetation. The scientific and applied effect of the research is creating the novel approach of development and practical use of intelligent technologies for agrotechnical monitoring by substantiating the methodological provisions for synthesis of structural and algorithmic organisation of the corresponding software and hardware solutions.

v2026.09.13