Skip to content

Add Magic Words embedding attack to Risks & mitigations#15

Open
WhymustIhaveaname wants to merge 1 commit into
bureado:mainfrom
SelfAdajoint:add-magic-words
Open

Add Magic Words embedding attack to Risks & mitigations#15
WhymustIhaveaname wants to merge 1 commit into
bureado:mainfrom
SelfAdajoint:add-magic-words

Conversation

@WhymustIhaveaname

Copy link
Copy Markdown

Adds a reference to "Jailbreaking LLMs' Safeguard with Universal Magic Words" under Risks & mitigations frameworks.

This paper demonstrates an embedding-layer attack that crafts universal suffixes to bypass RAG safeguards, relevant to agent runtime security where RAG pipelines filter tool inputs.

Paper: https://arxiv.org/abs/2501.18280

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant