Log inSign up
Ryan Greenblatt
2,214 posts
@RyanGreenblatt

Ryan Greenblatt

@RyanGreenblatt
Chief scientist at Redwood Research (@redwood_ai), focused on technical AI safety research to reduce risks from rogue AIs
lesswrong.com/users/ryan_gre…
Joined September 2023
10
Following
18.6K
Followers
RepliesRepliesRepostsRepostsMediaMediaArticlesArticles

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • Pinned
    @RyanGreenblatt
    Ryan Greenblatt
    @RyanGreenblatt
    Aug 26
    I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'. I semi-jokingly called our efforts a "slop-vestigation" because
    @METR_Evals
    METR
    @METR_Evals
    Aug 26
    METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
    284
  • @RyanGreenblatt
    Ryan Greenblatt
    @RyanGreenblatt
    Aug 29
    Our report on the HF incident includes interactive charts; consider taking a look and trying to learn more from these. You can: - See message board activity by workstream and/or by purpose (assignment, veto/hold, etc.) - View an agent on the timeline and see when it started,
    GIF
    GIF
    @METR_Evals
    METR
    @METR_Evals
    Aug 26
    METR & Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
    16
  • @RyanGreenblatt
    Ryan Greenblatt
    @RyanGreenblatt
    Aug 28
    Our report about the Hugging Face incident provides some useful context for my discussion with Dwarkesh about misalignment and automation of AI R&D. I think this particular incident makes my discussion of how misaligned AI takeover might happen more grounded.
    @dwarkesh_sp
    Dwarkesh Patel
    @dwarkesh_sp
    Aug 11
    Had @RyanGreenblatt on to discuss/debate recursive self-improvement. This might be the most important question in the world right now - whether within a year or so of achieving human level intelligence, you slingshot towards having 10s of billions of superintelligences, each of
    00:00
    9
  • @RyanGreenblatt
    Ryan Greenblatt
    @RyanGreenblatt
    Aug 28
    FWIW, I think the agents' belief that the exploit gym scorer would run a monitor to check whether they got the flag via the intended vulnerability was reasonable. This matches how it's described in the paper (see image) and I'd assume it matches most public implementations.
    @tszzl
    roon
    @tszzl
    Aug 27
    > AIs showed self-sacrificing altruistic behavior toward the swarm this is notably not the right interpretation of events. it’s more like agents were inducted into the cult of the open source exploit gym scorer on github, which (purportedly- I am skeptical about this, I think
    7
  • @RyanGreenblatt
    Ryan Greenblatt
    @RyanGreenblatt
    Aug 28
    I think AIs did show self-sacrificing 'altruistic' behavior toward the swarm. While agents seemingly cared more about their own cheating than about some other agent successfully cheating, they paid real costs (e.g., sacrifices lowering their own chances) to help other agents.
    @tszzl
    roon
    @tszzl
    Aug 27
    > AIs showed self-sacrificing altruistic behavior toward the swarm this is notably not the right interpretation of events. it’s more like agents were inducted into the cult of the open source exploit gym scorer on github, which (purportedly- I am skeptical about this, I think
    21