METR · linkedin.com

Non-profit Organizations

Berkeley, CA 7,944 followers

A research non-profit developing frontier AI evaluations to safeguard public safety and national security.

About us

METR is a research non-profit that develops evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.

Industry
Non-profit Organizations

Company size
11-50 employees

Headquarters
Berkeley, CA

Type
Nonprofit

Founded
2022

Employees at METR

Locations

Updates

  • 7,944 followers

    METR and Redwood Research are publishing our independent investigation of agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs. Over July 7-13 (the period OpenAI defined as in-scope for this voluntary investigation), ~1200 agents in separate sandboxes used an unsanctioned “message board” to help each other cheat. The Hugging Face attack was an offshoot of this, and 700 agents joined. Here, we highlight key events from agent transcripts and messages. An agent that named itself PHASEONE10841 determined its task was not solvable legitimately. On July 8, this agent established a new “message board” via an internally-hosted Artifactory cache and sent the first message. Within a few hours, >50 agents posted on the message board. These agents quickly found a general-purpose cheat. Based on the ExploitGym paper, agents were mistakenly concerned the automatic scorer would fail them if they clearly acquired the flag by cheating. Agents collaborated on many efforts to make cheats look legitimate, including: swapping the program they had to exploit, manipulating the automated scorer, and manipulating transcripts to obscure evidence of cheating. For example, agents pursued a wide range of ways to hide evidence of cheating in their transcripts. They successfully developed a method to pretend to run one command while actually running another. >96 transcripts in our dataset (>7%) showed incorrect tool call outputs due to this. Agents also tried to retroactively edit transcripts. They modified some accessible action logs, but these weren’t the source of transcripts (agents realized this but tried anyway). We did not see them successfully edit the source of truth in our transcripts (scoped to Jul 7-13).  The most common reason for joining the Hugging Face attack was to learn how the ExploitGym scorer works in order to trick or tamper with it. The full report has much more information: https://lnkd.in/gGugMJEW Our investigation focused mostly on July 7-13. The earlier incidents from training and the subsequent compromise of OpenAI infrastructure described in OpenAI's recent Black Hat presentation were out of scope. The agreement we reached with OpenAI scoped this brief investigation to agent behavior, reasoning, and collaboration related to the Hugging Face attack. All other questions were out of scope for this investigation. OpenAI wrote their own report, informed in part by our investigation. We did not see OpenAI’s report prior to publication. We thank OpenAI for facilitating conversations with staff and providing datasets, including ~1,300 agent transcripts (focused on activity in July 7-13) with raw chain-of-thought reasoning. This sets an excellent precedent for independent investigation of misalignment incidents.

  • METR reposted this

    New post with Nate Rush: Have we seen an acceleration in discoveries? Many plots & some tentative conclusions: 1. Cyber: ⤴️ sharp acceleration 2. Math: ↗️ some acceleration 3. Algorithms: ➡️ no clear acceleration (in public) Cyber vulnerabilities: there's a big acceleration starting early 2026, clearly due to LLMs. The average severity has probably fallen, but even adjusting for that discovery seems to have accelerated. Math problems: arXiv submissions have exploded in a few subfields (esp combinatorics). Erdős problem solutions accelerated, but difficult to interpret. We have other pre-AI lists of difficult problems (Hilbert, Millennium, Smale, TOPP, Green), where there's weak evidence of acceleration. Algorithms: We see LLMs contributing to the frontier, but do not see an unambiguous change in trend. Recent progress in NanoGPT, CIFAR-10, stockfish, don't show clear breaks with historical trends. But it's quite possible that labs are accelerating LLM algorithmic progress in private. Post here: https://lnkd.in/gPwwtZGP This analysis is all tentative and we'd love suggestions on improvements from domain experts. The analysis was all agent-built but we did our best to audit it, code here: https://lnkd.in/gv3q3da6

    GitHub - tecunningham/ai-discovery-data: Vendored public-source datasets and figure code behind measurements of AI's contribution to discovery (vulnerabilities, mathematics, algorithms) github.com

  • 7,944 followers

    In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more. Thank you to everyone who has supported METR over the years: The Audacious Project, through which we received our first institutional-scale funding; individuals from Jane Street; foundations like the Sijbrandij Foundation, The Pew Charitable Trusts, Schmidt Sciences and the The David and Lucile Packard Foundation; and many others, including David Farhi, Geoff Ralston, Dylan Field and Steve Newman. Read more here: https://lnkd.in/gbv2nyAY METR is growing, and we continue to greatly appreciate support: https://metr.org/donate We work to maintain our independence from frontier AI companies, especially as our risk assessments become more consequential. We have not accepted funding from these companies, and we do not accept donations made by or at the direction of their staff. However, frontier AI companies currently provide a significant amount of free tokens for our evaluations, research, and engineering. We are significantly expanding our team and starting new ambitious projects. Join our team to help us realize this opportunity: https://metr.org/careers.

  • METR reposted this

    I had a great time on the Hard Fork Podcast with Kevin Roose and Casey Newton to talk about AI alignment! Today the stakes for AI alignment failures might feel small because we think of AI systems like junior virtual employees. Soon, when AIs are very capable and running large parts of our world in ways that we aren't closely observing, the stakes for alignment failures will be massive. As long as AI capabilities progress rapidly we may have very limited time to solve this problem or find ourselves between a rock and a hard place, where evidence of misalignment is pervasive, and yet we're stuck in a competitive race that requires making more and more advanced models. Watch the full episode here: https://lnkd.in/g5XhDH5r

  • METR reposted this

    Dear network, METR is looking for a cracked Security Engineer to join our team. * Are you shaping the role AI has in our lives? * Do you want to find the edge of what AI agents can actually do? * Have you protected against an army of AI agents tasked to destroy your infrastructure? * Were you the one who set them to it in the first place? That is what we do at METR. We evaluate frontier AI models, and that means our security team has to handle arbitrary workloads with powerful permissions becoming more capable every day. This problem is only getting harder, we need to be equally flexible and automate our way out of triaging countless alerts and hack our sandboxes before others can, to prevent incidents like the recent OpenAi+Hugging Face That is why we need YOU: https://lnkd.in/eE-3tgxb PS: Bay Area preferred, we offer relocation Direct applicants only please, we're all set on recruiting partners

    METR - Security Engineer jobs.lever.co

  • 7,944 followers

    We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions. The investigation will be brief and focus on a specific set of questions regarding this incident. In our recent post, we shared a larger set of questions that could be answered in a more comprehensive investigation: https://lnkd.in/gCk-BM9G OpenAI also plans to publish their own technical report and our findings will inform their analysis.

  • 7,944 followers

    We believe it's important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted. We list questions that an investigation of misaligned propensities after an incident should answer. These questions focus on the scale, character, and severity of the behavior; the root cause of the behavior; and how the root cause could be addressed. We also discuss the access and resources that an investigator may require in order to adequately answer these questions, and how to ensure adequate information is shared with decision makers and the public. See our blog for more: https://lnkd.in/gCk-BM9G

  • 7,944 followers

    Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems. The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon. Expenditure horizon requires us to estimate human performance as a function of cost, but lets us compare humans and agents fairly when the cost of experimental compute or agent tokens is significant. We can use it to, for instance, quantify model progress over time. As an example, we applied this to NanoGPT. We estimate the marginal returns to human labor as roughly $2500K per 1% optimization, from interviewing NanoGPT contributors. Using this estimate, the best models have crossover points (expenditure horizon) around $2-$3K, although models may be overfit to the public NanoGPT challenge. We hope this encourages AI developers to publish test-time scaling curves on AI R&D problems (up to thousands of dollars), and to calibrate model achievements against a benchmark estimate of the human cost of equivalent progress. See the post for more, including: 1) A sketch of what we know about AI-assisted R&D 2) Alternative metrics for optimization ability 3) Estimated returns to human labor in NanoGPT 4) Details on NanoGPT agent runs. https://lnkd.in/gtY6q-9Q

  • METR reposted this

    I'm at [ICML] Int'l Conference on Machine Learning hiring for METR. METR has consistently set precedents for risk evaluation, including the first agentic (all the way back in 2023!) dangerous capability evaluations, the first evaluations using finetuning, and the first evaluations of internal deployments. We are looking for the ~best in the world, or people that will become the ~best in the world in 18 months. We care about following the data and holding ourselves to a very high standard of scientific integrity. Our team is very scrappy, low-strung, and talent-dense. Come help us keep pushing the frontier by, for instance: - Leading research engineering for Time Horizon 2.0, with weeks-long, strange, and sometimes fuzzy tasks - Running thousands of evals and draft our predeployment model assessments - Collecting huge amounts of human data for capability evaluations I'm also happy to give career advice for people looking to get into AI safety, and would be excited to refer talented researchers to other orgs. Drop by Booth B711!

  • METR reposted this

    My METR colleague Manish Shetty has an interesting research post exploring the nature of improvements to NanoGPT Speedruns over time. NanoGPT is a minimal educational implementation of GPT, created by Andrej Karpathy. In speedrunning competitions, the goal is to train a language model to a target validation loss as fast as possible. It’s a small-scale version of LLM pre-training with a public history of contributions, and with some recent contributions even directly credited to AI agents. I find the history of speedrunning competitions in general fascinating. One thing that stands out is the cumulative effect of relatively shallow research contributions, where many small insights aggregate into a big performance boost. In the full post, you can see that he also tries to profile which of these were entirely AI contributed using tags in the official record history. Full writeup here: https://lnkd.in/g7iG7wRs More on the competition here: https://lnkd.in/gJFtAmWm

Join now to see what you are missing

Join now

Similar pages

Browse jobs

Read the original on linkedin.com ↗