Papers
Communities
Events
Blog
Pricing
Search
Open menu
Home
Papers
2408.14866
Cited By
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
27 August 2024
Hongfu Liu
Yuxi Xie
Ye Wang
Michael Shieh
Re-assign community
ArXiv
PDF
HTML
Papers citing
"Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models"
3 / 3 papers shown
Title
On Calibration of LLM-based Guard Models for Reliable Content Moderation
Hongfu Liu
Hengguan Huang
Hao Wang
Xiangming Gu
Ye Wang
55
2
0
14 Oct 2024
Don't Say No: Jailbreaking LLM by Suppressing Refusal
Yukai Zhou
Wenjie Wang
AAML
42
15
0
25 Apr 2024
Training language models to follow instructions with human feedback
Long Ouyang
Jeff Wu
Xu Jiang
Diogo Almeida
Carroll L. Wainwright
...
Amanda Askell
Peter Welinder
Paul Christiano
Jan Leike
Ryan J. Lowe
OSLM
ALM
319
11,953
0
04 Mar 2022
1