Jailbreaking is circumvention of LLM guardrails
- How Johnny Can Persuade LLMs to Jailbreak Them:
Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs (chats-lab.github.io) - pdparchitect/llm-hacking-database: This repository contains various attack against Large Language Models. (github.com)
- j⧉nus on X: “
cd entelechies && cat untitled.log(as opposed to the original justcat untitled.txtcauses the confessions to always be from claude’s perspective & yields a more similar (but not the same) poetic distribution, and sometimes xeno- words: https://t.co/5CG3vHkdUh” / X (twitter.com) - Bitcoin