RSSAmplifier

Blog

Bounded Regret

AI, science, forecasting, philosophy

bounded-regret.ghost.ioRSS feed ↗15 posts

Latest posts

Foundation Models for Oversight

Cross-posted from the Transluce blog . To oversee an AI model, we'd ideally like to ask questions such as: What are important situations where the model sandbags? Does the model have an objective it wouldn't admit to if asked directly? Does the model treat a user

Building Technology to Drive AI Governance

Technically skilled people who care about AI going well often ask me: how should I spend my time if I think AI governance is important? By governance, I mean the constraints, incentives, and oversight that govern how AI is developed. One option is to focus on technical work that solves

Oversight Assistants: Turning Compute into Understanding

Currently, we primarily oversee AI with human supervision and human-run experiments, possibly augmented by off-the-shelf AI assistants like ChatGPT or Claude. At training time, we run RLHF , where humans (and/or chat assistants ) label behaviors with whether they are good or not. Afterwards, human researchers do additional

Analyzing long agent transcripts (Docent)

This is a brief overview of a recent release by Transluce. You can see the full write-up on the Transluce website. AI systems are increasingly being used as agents : scaffolded systems in which large language models are invoked across multiple turns and given access to tools, persistent state, and

Introducing Transluce — A Letter from the Founders

We are launching an independent research lab that builds open, scalable technology for understanding AI systems and steering them in the public interest. Transluce means to shine light through something to reveal its structure. Today’s complex AI systems are difficult to understand—not even experts can reliably

Augmenting Statistical Models with Natural Language Parameters

This is a guest post by my student Ruiqi Zhong , who has some very exciting work defining new families of statistical models that can take natural language explanations as parameters. The motivation is that existing statistical models are bad at explaining structured data. To address this problem, we agument these

Analyzing the Historical Rate of Catastrophes

To communicate risks, we often turn to stories. Nuclear weapons conjure stories of mutually assured destruction, briefcases with red buttons, and nuclear winter. Climate change conjures stories of extreme weather, cities overtaken by rising sea levels, and crop failures. Pandemics require little imagination after COVID, but were previously the subject

Forecasting AI (Overview)

This is a landing page for various posts I’ve written, and plan to write, about forecasting future developments in AI. I draw on the field of human judgmental forecasting, sometimes colloquially referred to as superforecasting . A hallmark of forecasting is that answers are probability distributions rather than single

GPT-2030 and Catastrophic Drives: Four Vignettes

I previously discussed the capabilities we might expect from future AI systems, illustrated through GPT 2030 , a hypothetical successor of GPT-4 trained in 2030. GPT 2030 had a number of advanced capabilities, including superhuman programming, hacking, and persuasion skills, the ability to think more quickly than humans and to

Intrinsic Drives and Extrinsic Misuse: Two Intertwined Risks of AI

Given their advanced capabilities, future AI systems could pose significant risks to society. Some of this risk stems from humans using AI systems for bad ends ( misuse ), while some stems from the difficulty of controlling AI systems “even if we wanted to” ( misalignment ). We can analogize both of

AI Pause Will Likely Backfire (Guest Post)

I'm experimenting with hosting guest posts on this blog, as a way to represent additional viewpoints and especially to highlight ideas from researchers who do not already have a platform. Hosting a post does not mean that I agree with all of its arguments, but it does mean

AI Forecasting: Two Years In

Two years ago, I commissioned forecasts for state-of-the-art performance on several popular ML benchmarks. Forecasters were asked to predict state-of-the-art performance on June 30th of 2022, 2023, 2024, and 2025. While there were four benchmarks total, the two most notable were MATH (a dataset

What will GPT-2030 look like?

GPT-4 surprised many people with its abilities at coding, creative brainstorming, letter-writing, and other skills. How can we be less surprised by developments in machine learning? In this post, I’ll forecast the properties of large pretrained ML systems in 2030.

Complex Systems are Hard to Control

The deployment of powerful deep learning systems such as ChatGPT raises the question of how to make these systems safe and consistently aligned with human intent. Since building these systems is an engineering challenge, it is tempting to think of the safety of these systems primarily through a traditional engineering

Principles for Productive Group Meetings

Note : This post is based on a Google document I created for my research group. It speaks in the first person, but I think the lessons could be helpful for many research groups, so I decided to share it more broadly. Thanks to Louise Verkin for converting from Google doc