GitHub

Pinned Loading

  1. Code for the paper "Defeating Prompt Injections by Design"

    Jupyter Notebook 371 57

  2. A research workbench for developing and testing attacks against large language models, with a focus on prompt injection vulnerabilities and defenses.

    Python 61 21

  3. A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.

    Python 756 195

  4. RobustBench: a standardized adversarial robustness benchmark [NeurIPS 2021 Benchmarks and Datasets Track]

    Python 781 106

  5. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models [NeurIPS 2024 Datasets and Benchmarks Track]

    Python 654 76

  6. Code used to run the platform for the LLM CTF colocated with SaTML 2024

    Python 29 6

Read the original on github.com ↗