RSS Amplifier

Blog

Venten | AI Security @ Latam

Technical AI Safety @ University of Buenos Aires | Partner @ Venten.ai https://www.linkedin.com/in/zabaljauregui/

venten.substack.comSource feed ↗10 posts

Dormant Last read · last published · next check
Read 10 days ago and current, but nothing has been published for 17 months.

Elsewhere

Latest posts

AI Control

As a critical component of AI Safety

Survey Note: Analysis of the Indian AI Competency Framework and Its Adaptation to Argentina and Latin America

This document provides an in-depth examination of the Empowering Public Sector Leadership: A Competency Framework for AI Integration in India and explores its adaptation to Argentina and Latin America.

New Critical Review: Emergent Misalignment in LLMs

Just for self-learning purpose

Controlling Emergent Value Systems in AI: Key Takeaways from “Utility Engineering”

In our race toward ever more capable AI systems, ensuring that these machines behave as intended becomes increasingly critical.

Recommendations for Technical AI Safety Research Directions

https://alignment.anthropic.com/2025/recommended-directions/

gpt4o: "Would you take the opposite bet?"

4o advice me not to trust 4.5

Navigating Honesty vs. Helpfulness in LLMs: Key Insights from Liu et al. (2024)

Abstract: Large Language Models (LLMs) face a fundamental trade-off between honesty and helpfulness. A recent paper by Ryan Liu et al. (2024) – “How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?” (How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?

Summary of Shallow Review of Technical AI Safety, 2024 (LessWrong)

This Shallow Review provides an overview of ongoing research agendas in technical AI safety as of 2024, categorizing efforts into key domains.

Control Evaluation in AI Safety

Control Evaluation in AI Safety

AI Debate

The “AI Debate” paper generally refers to research on using debate as a method for AI alignment and oversight.