
AI Control
As a critical component of AI Safety
Technical AI Safety @ University of Buenos Aires | Partner @ Venten.ai https://www.linkedin.com/in/zabaljauregui/
Dormant Last read · last published · next check
Read 10 days ago and current, but nothing has been published for 17 months.

As a critical component of AI Safety

This document provides an in-depth examination of the Empowering Public Sector Leadership: A Competency Framework for AI Integration in India and explores its adaptation to Argentina and Latin America.

Just for self-learning purpose

In our race toward ever more capable AI systems, ensuring that these machines behave as intended becomes increasingly critical.

https://alignment.anthropic.com/2025/recommended-directions/

4o advice me not to trust 4.5

Abstract: Large Language Models (LLMs) face a fundamental trade-off between honesty and helpfulness. A recent paper by Ryan Liu et al. (2024) – “How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?” (How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?

This Shallow Review provides an overview of ongoing research agendas in technical AI safety as of 2024, categorizing efforts into key domains.

Control Evaluation in AI Safety

The “AI Debate” paper generally refers to research on using debate as a method for AI alignment and oversight.