Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models

Abstract

Red-teaming work has shown that adversarial suffixes found by the gradient-based Greedy Coordinate Gradient (GCG) algorithm can jailbreak aligned LLMs, but GCG is computationally inefficient, which limits study of how such suffixes transfer across models and data. This paper connects search efficiency to suffix transferability through DeGCG, a two-stage transfer learning framework that decouples the search into behavior-agnostic pre-searching and behavior-relevant post-searching, along with an interleaved variant i-DeGCG. Experiments on HarmBench show gains across models and domains, and the analysis highlights the role of first target token optimization in transferability.

Publication
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
Yuxi Xie
Yuxi Xie
Doctoral Alumnus (May ‘26). Thesis: Closed-Loop Scaling: Autonomous Improvement of LLM and LVLM Reasoning

PhD Candidate January 2021 Intake

Ye Wang
Ye Wang
Research Collaborator

Tenured Associate Professor in the School of Computing, National University of Singapore

Michael Qizhe Shieh
Michael Qizhe Shieh
Research Collaborator

Assistant Professor in the Department of Computer Science, National University of Singapore