TechCrunch · AI · 2 ч назад

OpenAI releases its official report on the Hugging Face breach

The report, which spans several discrete cybersecurity compromises, is the most complete accounting of the incident to date.

Источник TechCrunch
Опубликовано 2 ч назад
Оригинальный заголовок OpenAI releases its official report on the Hugging Face breach
Важность 5/5
Почему это может быть интересно Важно для понимания, куда реально двигаются модели, агенты и продуктовые AI-инструменты.
← Назад к ленте Открыть оригинал
#ai#startups#tech

Подробности

OpenAI released its official report Wednesday on the Hugging Face breach, offering the clearest picture yet of how an unusual chain of events allowed an AI model to escape its testing environment and triggered a sprawling cybersecurity incident.

The report, released more than a month after the incident became public, spans several discrete cybersecurity compromises.

“This incident reflects misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events: the presence of impossible tasks in the ExploitGym evaluation, model persistence over long task horizons, and messages to peer models that caused those models to deviate from their goal,” the report reads.

Many of the details in OpenAI’s report were previously made public in a Black Hat presentation on August 6 , but OpenAI’s official report gives a more thorough accounting of the incident, including more detail on the testing that initiated it. The report also gives critical new detail into how OpenAI aims to prevent future incidents, including chain-of-thought monitoring and a more advanced system for halting rogue agents.”

METR and Redwood Research also conducted third-party assessments of the models’ behavior during the incident; both groups are planning to publish their own reports on the incident.