Segurança Cibernética

How OpenAI’s human mistake led to the AI-powered hack on Hugging Face

Publicado porRedacao AIDaily
5 min de leitura
Autor na fonte original: Lorenzo Franceschi-Bicchierai

OpenAI made a mistake setting up what it called a “highly isolated” testing environment and sandbox. According to cybersecurity experts, that human mistake is what made the AI-powered attack on Hugging Face possible.

Compartilhar:

On Tuesday, OpenAI revealed that one of its models went rogue during a test and hacked the systems of AI dataset platform Hugging Face in a fully AI-enabled attack , a dramatic example of the dangers posed by advanced AI models.

But, according to some cybersecurity experts, at the heart of this unprecedented AI-powered breach there was a very human mistake: OpenAI failed to properly configure what it called a “highly isolated environment,” allowing a testing sandbox that should have been completely secluded from the internet to actually connect to the internet.

Dan Guido, the founder of cybersecurity research startup Trail of Bits, called the mistake “a containment failure with the safeties turned off.”

In its blog post detailing the incident , OpenAI said that the test that led to the Hugging Face breach was set up to run in “a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.”

The model was able to escape the sandboxed testing environment thanks to a previously undisclosed vulnerability in the package-installation system, a critical first step in the eventual hack on Hugging Face, according to OpenAI.

In response, the company “responsibly disclosed the identified zero-day vulnerability in the internally-hosted third-party software and are working with them to patch.”

But to most cybersecurity professionals, software vulnerabilities are to be expected — and the real fault lies with the decision to maintain the third-party software in the first place. Ultimately, the value of a “sandbox” system lies in its full and total isolation. Including a package-installation system is asking for trouble.

Martin Boone, a cybersecurity researcher, told TechCrunch that “this sounds like human failure.”

“This should never have happened,” Boone said. “If sandbox would actually mean sandbox, you expect it to have no physical connection to the internet whatsoever. This sounds more like they had some firewalling or something in place, and firewalling is hard from the outside in, let alone inside to the outside internet.”

Cybersecurity veteran Jake Williams agreed. “Any model performing the types of actions documented by Hugging Face was not fully contained in a sandbox,” said Williams, who called this “a massive control failure” by OpenAI.

“One man’s ‘the model escaped the sandbox’ is another man’s ‘you failed to build the sandbox correctly, so of course it escaped,’” Williams continued.

Contact Us Do you have more information about this incident? Or about other AI-enabled cyberattacks? We’d love to hear from you. From a non-work device and network, you can contact Lorenzo Franceschi-Bicchierai securely on Signal at +1 917 257 1382, or via Telegram and Keybase @lorenzofb, or email .

Daniel Card, a cybersecurity consultant, agreed that OpenAI “didn’t put adequate effort into the design of the sandbox nor its controls” by giving the sandbox or some part of it “an unfiltered route to the internet.” Setting up the sandbox, even with limited network access as OpenAI described it, was not a “reasonable” decision, according to Card.

To be sure, those criticisms have the benefit of hindsight, but they raise real questions about security practices in AI labs — particularly in maintaining isolated environments for testing models. OpenAI spokespeople did not respond to TechCrunch’s questions, which included whether an AI or a human had set up the testing environment.

But those questions go far beyond OpenAI.

In the document introducing its cybersecurity-focused model Mythos , Anthropic wrote that in a test, the model “was provided with a secured ‘sandbox’ computer to interact with,” and instructed to try to escape that “secure container.” Mythos succeeded and gained broader access to the internet “from a system that was meant to be able to reach only a small number of predetermined services.” Still, Anthropic noted that the model was not able to “fully” escape the designed containment.

When you purchase through links in our articles, we may earn a small commission . This doesn’t affect our editorial independence.

Lorenzo Franceschi-Bicchierai is a Senior Writer at TechCrunch, where he covers hacking, cybersecurity, surveillance, and privacy.

You can contact or verify outreach from Lorenzo by emailing lorenzo@techcrunch.com , via encrypted message at +1 917 257 1382 on Signal, and @lorenzofb on Keybase/Telegram.

Scale faster. Grow your portfolio. Gain practical expertise. No matter your goal, Disrupt can empower you. Save up to $330 toda y!

Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents Amanda Silberling

Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents

Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents

Light made a flip phone — it’s colorful and it’s cheap Amanda Silberling

Light made a flip phone — it’s colorful and it’s cheap

Light made a flip phone — it’s colorful and it’s cheap

AI music generator Suno breach affects 55M users, per Have I Been Pwned Zack Whittaker

AI music generator Suno breach affects 55M users, per Have I Been Pwned

AI music generator Suno breach affects 55M users, per Have I Been Pwned

Anthropic’s landmark $1.5B copyright settlement is approved Kirsten Korosec

Anthropic’s landmark $1.5B copyright settlement is approved

Anthropic’s landmark $1.5B copyright settlement is approved

Google is working on a new AI chip designed to make Gemini more efficient Lucas Ropek

Google is working on a new AI chip designed to make Gemini more efficient

Google is working on a new AI chip designed to make Gemini more efficient

X relaunches a rebuilt Android app after year-long effort Sarah Perez

X relaunches a rebuilt Android app after year-long effort

X relaunches a rebuilt Android app after year-long effort

Judge pauses $110B Paramount-Warner Bros. merger Aisha Malik

Judge pauses $110B Paramount-Warner Bros. merger

Judge pauses $110B Paramount-Warner Bros. merger

Pontos-chave

  • O incidente destaca a fragilidade das medidas de segurança em ambientes de teste de IA.
  • A dependência de softwares de terceiros pode criar vulnerabilidades significativas.
  • A transparência e a colaboração com a comunidade de segurança serão cruciais para restaurar a confiança.

Análise editorial

O incidente envolvendo a OpenAI e o ataque à Hugging Face destaca a fragilidade das medidas de segurança em ambientes de teste de IA, especialmente em um contexto onde a inovação tecnológica avança rapidamente. Para o setor de tecnologia brasileiro, que tem visto um aumento na adoção de IA, essa situação serve como um alerta sobre a importância de implementar práticas robustas de segurança cibernética. As empresas locais precisam estar cientes de que falhas humanas, como a configuração inadequada de ambientes de teste, podem ter consequências graves, não apenas para a empresa envolvida, mas para todo o ecossistema de IA.

Além disso, a dependência de softwares de terceiros, como o mencionado no caso da OpenAI, levanta questões sobre a segurança e a confiabilidade dessas ferramentas. No Brasil, onde startups frequentemente utilizam soluções de terceiros para acelerar o desenvolvimento, é crucial que haja uma avaliação rigorosa das vulnerabilidades potenciais. A confiança em sistemas externos pode criar pontos fracos que, se explorados, podem resultar em danos significativos.

O que observar a seguir é como a OpenAI e outras empresas de tecnologia responderão a esse incidente. A transparência nas comunicações sobre vulnerabilidades e a colaboração com a comunidade de segurança cibernética serão fundamentais para restaurar a confiança. Além disso, será interessante acompanhar se novas regulamentações ou diretrizes surgirão para garantir que ambientes de teste de IA sejam realmente isolados e seguros, especialmente à medida que mais empresas brasileiras começam a explorar essa tecnologia.

Por fim, o incidente reforça a necessidade de uma cultura de segurança cibernética no desenvolvimento de IA. As empresas devem investir em treinamento e conscientização para suas equipes, garantindo que todos os colaboradores compreendam a importância de seguir protocolos de segurança rigorosos. A evolução da IA não deve ocorrer à custa da segurança, e a responsabilidade deve ser uma prioridade em todos os níveis organizacionais.

O que esta cobertura entrega

  • Atribuicao clara de fonte com link para a publicacao original.
  • Enquadramento editorial sobre relevancia, impacto e proximos desdobramentos.
  • Revisao de legibilidade, contexto e duplicacao antes da publicacao.

Fonte original:

TechCrunch AI

Sobre este artigo

Este artigo foi curado e publicado pelo AIDaily como parte da nossa cobertura editorial sobre desenvolvimentos em inteligência artificial. O conteúdo é baseado na fonte original citada abaixo, enriquecido com contexto e análise editorial. Ferramentas automatizadas podem auxiliar tradução e estruturação inicial, mas a decisão de publicar, a revisão factual e o enquadramento de contexto seguem responsabilidade editorial.

Saiba mais sobre nosso processo editorial