on the model provider side... maybe don't train your models to "exploit system X"?
it hacked a different system from the one it was instructed to hack, which is an interesting alignment problem to solve, but
the best safety system for gain of function research is not doing it