RSS Amplifier

Lukas Finnveden · Jan 3, 2024

Project ideas for making transformative AI go well, other than by working on alignment

0
Sign in to vote or save

Lukas Finnveden · Lukas Finnveden

This series of posts contains lists of projects that it could be valuable for someone to work on. The unifying theme is that they are projects that:

  • Would be especially valuable if transformative AI is coming in the next 10 years or so.

  • Are not primarily about controlling AI or aligning AI to human intentions.1

    • Most of the projects would be valuable even if we were guaranteed to get aligned AI.

    • Some of the projects would be especially valuable if we were inevitably going to get misaligned AI.

The posts contain some discussion of how important it is to work on these topics, but not a lot. For previous discussion (especially: discussing the objection “Why not leave these issues to future AI systems?”), you can see the section How ITN are these issues? from my previous memo on some neglected topics.

The lists are definitely not exhaustive. Failure to include an idea doesn’t necessarily mean I wouldn’t like it. (Similarly, although I’ve made some attempts to link to previous writings when appropriate, I’m sure to have missed a lot of good previous content.)

There’s a lot of variation in how sketched out the projects are. Most of the projects just have some informal notes and would require more thought before someone could start executing. If you're potentially interested in working on any of them and you could benefit from more discussion, I’d be excited if you reached out to me!2

There’s also a lot of variation in skills needed for the projects. If you’re looking for projects that are especially suited to your talents, you can search the posts3 for any of the following tags (including brackets):

[ML]   [Empirical research]   [Philosophical/conceptual]   [survey/interview]   [Advocacy]   [Governance]   [Writing]   [Forecasting]

The projects are organized into the following categories (which are in separate posts). Feel free to skip to whatever you’re most interested in.

(If you want to comment on any of the posts in this series, you could do so either here, at the EA forum, or on LessWrong.)

Few of the ideas in these posts are original to me. I’ve benefited from conversations with many people. Nevertheless, all views are my own.

For some projects, I credit someone who especially contributed to my understanding of the idea. If I do, that doesn’t mean they have read or agree with how I present the idea (I may well have distorted it beyond recognition). If I don’t, I’m still likely to have drawn heavily on discussion with others, and I apologize for any failure to assign appropriate credit.

For general comments and discussion, thanks to Joseph Carlsmith, Paul Christiano, Jesse Clifton, Owen Cotton-Barrat, Daniel Kokotajlo, Linh Chi Nguyen, Fin Moorhouse, Caspar Oesterheld, and Carl Shulman.

Here’s a list with all the project ideas from the other posts. (Sorry it’s not hyper-linked.) Unless you’re looking for something specific, I suggest jumping into the first post instead of reading this.

Project ideas: Governance during explosive technological growth

  • Investigate and publicly make the case for/against explosive growth being likely and risky [Forecasting] [Empirical research] [Philosophical/conceptual] [Writing]

  • Painting a picture of a great outcome [Forecasting] [Philosophical/conceptual] [Governance]

  • Policy-analysis of issues that could come up with explosive technological growth [Governance] [Forecasting] [Philosophical/conceptual]

    • Address vulnerable world hypothesis with minimal costs

    • How to handle brinkmanship/threats?

    • Avoiding AI-assisted human coups

    • Governance issues raised by digital minds

  • Norms/proposals for how to navigate an intelligence explosion [Governance] [Forecasting] [Philosophical/conceptual]

    • No “first strike” intelligence explosion

    • Never go faster than X?

    • Concrete decision-making proposals

    • Technical proposals for slowing down / coordinating

    • Dubiously enforceable promises

    • Technical proposals for aggregating preferences

  • Decrease the power of bad actors

    • Avoid malevolent individuals getting power within key organizations [Governance]

    • Prevent dangerous external individuals/organizations from having access to AI [Governance] [ML]

    • Accelerate good actors

  • Analyze: What tech could change the landscape? [Forecasting] [Philosophical/conceptual] [Governance]

  • Big list of questions that labs should have answers for [Philosophical/conceptual] [Forecasting] [Governance]

Project ideas: Epistemics

  • Why AI matters for epistemics

  • Why working on this could be urgent

  • Categories of projects

  • Differential technology development [ML] [Forecasting] [Philosophical/conceptual]

    • Important subject areas

    • Methodologies

    • Related/previous work.

  • Get AI to be used & (appropriately) trusted

    • Develop technical proposals for how to train models in a transparently trustworthy way [ML] [Governance]

    • Survey groups on what they would find convincing [survey/interview]

    • Create good organizations or tools [ML] [Empirical research] [Governance]

    • Examples of organizations or products

    • Investigate and publicly make the case for why/when we should trust AI about important issues [Writing] [Philosophical/conceptual] [Advocacy] [Forecasting]

    • Developing standards or certification approaches [ML] [Governance]

  • Develop & advocate for legislation against bad persuasion [Governance] [Advocacy]

Project ideas: Sentience and rights of digital minds

  • Develop & advocate for lab policies [ML] [Governance] [Advocacy] [Writing] [Philosophical/conceptual]

    • Create an RSP-style set of commitments for what evaluations to run and how to respond to them

    • Policies that don’t require sophisticated information about AI preferences/experiences

    • Preserving models for later reconstruction

    • Deploy in “easier” circumstances than trained in

    • Reduce extremely out of distribution (OOD) inputs

    • Train or prompt for happy characters

    • Committing resources to research on AI welfare and rights

    • Learning more about AI preferences

    • Credible offers

    • Talking via internals

    • Training for honest self-reports

    • Clues from AI generalization

    • Interpretability

    • Interventions that rely on understanding AI preferences

    • Offer an alternative to working (exit, sleep, or retirement)

    • Commitment to pay AI systems

    • Tell the world

    • Train AI systems that suffer less and have fewer preferences that are hard to satisfy

  • Investigate and publicly make the case for/against near-term AI sentience or rights [Philosophical/conceptual] [Writing]

  • Study/survey what people (will) think about AI sentience/rights [survey/interview]

  • Develop candidate regulation [Governance] [Forecasting]

  • Avoid inconvenient large-scale preferences [Philosophical/conceptual]

  • Advocating for statements about digital minds [Governance] [Advocacy] [Writing]

Project ideas: Backup plans & Cooperative AI

  • Backup plans for misaligned AI

    • What properties would we prefer misaligned AIs to have? [Philosophical/conceptual] [Forecasting]

      • Making misaligned AI have better interactions with other actors

      • AIs that we may have moral or decision-theoretic reasons to empower

      • Making misaligned AI positively inclined toward us

    • Studying generalization & AI personalities to find easily-influenceable properties [ML]

    • Theoretical reasoning about generalization [ML] [Philosophical/conceptual]

  • Cooperative AI

    • Implementing surrogate goals / safe Pareto improvements [ML] [Philosophical/conceptual] [Governance]

    • AI-assisted negotiation [ML] [Philosophical/conceptual]

    • Implications of acausal decision theory [Philosophical/conceptual]

1

Nor are they primarily about reducing risks from engineered pandemics.

2

My email is [last name].[first name]@gmail.com

3

Or the table of content below.

Read the original on lukasfinnveden.substack.com

Comments

Nothing yet. Say the first thing.

    Sign in to join the conversation.