GitHub

View max-andr's full-sized avatar

🚀

Maksym Andriushchenko max-andr

🚀

Block or report max-andr

Pinned Loading

  1. Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours

    Python 518 59

  2. OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents [NeurIPS 2025 Spotlight]

    Jupyter Notebook 69 6

  3. Does Refusal Training in LLMs Generalize to the Past Tense? [ICLR 2025]

    Python 78 12

  4. Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks [ICLR 2025]

    Shell 393 47

  5. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models [NeurIPS 2024 Datasets and Benchmarks Track]

    Python 654 76

  6. RobustBench: a standardized adversarial robustness benchmark [NeurIPS 2021 Benchmarks and Datasets Track]

    Python 781 106

Read the original on github.com ↗