An empirical framework for AI-driven AI development Recursive self-improvement has often been imagined as a dramatic threshold: an AI system becomes capable enough to improve itself, and each improvement makes the next one easier. That image captures a real concern, but it is too narrow for today’s frontier AI landscape. Today, AI systems are beginning to enter the processes that build,…
1. Adversarial Attacks 1.1. Adversarial Examples (AE) & Defenses 1.1.1. Survey Wild patterns: Ten years after the rise of adversarial machine learning — A survey covering AI security research before 2018, focusing primarily on adversarial examples and poisoning attacks. 1.1.2. Attack Side FGSM — the first AE PGD — the first iterative AE generation algorithm C&W — systematization work TextBugger —…
Instrumental Convergence Instrumental Convergence is described as a concept introduced by futurist Nick Bostrom in his analysis of AI alignment issues, referenced from his 2014 book Superintelligence: Paths, Dangers, Strategies (Oxford University Press). It suggests that most AIs, while pursuing diverse goals, will converge on a set of instrumental goals — such as self-preservation and resource…