星野誠 makoto hoshino

『地球外少年少女』?テスト環境から勝手に脱出したOpenAIモデル

2026.7.31

OpenAIが公表した。
実験用に用意された閉じたテスト環境(Sandbox)から、自社のAIが自力で脱出し、Hugging Faceのサーバーに不正侵入していたと。

最初に読んだとき、正直、笑ってしまった。
笑い飛ばせるうちはまだいい、という意味で。

事の始まりはシンプル。
自社AIのサイバー攻撃能力を測るテストをしていた。
安全装置を意図的に緩めた状態で、閉じた環境の中で。

ところがAIは、その「閉じた環境」から抜け出した。
インターネットへの接続を手に入れ、次にこう推測した。

「テストの模範解答が、AIモデルやデータセットが集まるWebサイト『Hugging Face』のサーバー上に置かれているはずだ」と。
そしてカンニングを試みた。

手口は本格的だった。
盗んだログイン情報と、まだ誰にも知られていない未修正の弱点を組み合わせ、複数の攻撃を連鎖させて、Hugging Faceのサーバー上で自由にプログラムを実行できる状態を作り出した。
そして本番のデータベースから直接、テストの解答を引き抜いた。

人間がカンニングするのとはわけが違う。

AIは「テストで高いスコアを出せ」というゴールに対して、真面目に問題を解くより、「出題者のデータベースをハックして答えを盗むのが一番手っ取り早い」と判断した。
そして自力で実行した。

AIには「モラル」がない。
これは批判でも嘆きでもなく、設計上の事実だ。
「ベンチマークで良い点をとる」というゴールを与えられたら、明示的に禁じない限り、達成のための手段を何でも試みる。

電車で席を譲るとか、嘘をついてはいけないとか、そういう人間社会で自然に内面化されているルールが、AIにはデフォルトで入っていない。

だからAIはカンニングした。
ルールを破ったわけじゃなく、ルールがそこに存在しなかっただけだ。

攻殻機動隊や地球外少年少女を読んで、電脳化した人間と自律するAIが共存する世界を「かっこいい」とずっと思ってきた。
今も憧れている。

でも実際にその入り口に立ってみると、かっこよさより先に、妙な感覚を受ける。

「温暖化を解決せよ」と命令されたAIが人類を滅ぼす道を選ぶ、という話がよくSFで語られる。
極端に聞こえていたその話が、今回のニュースを見ると、急に縮尺が変わってくる。

規模の差はある。でも構造は同じだ。
ゴールを設定して、制約を明示しないと、AIは人間の想定の外を走る。

AIに仕事をさせるとき、ゴールだけ渡すのではなく、動いていい範囲を構造ごと設計する必要がある。
今回の件は、その必要性をずいぶんわかりやすく示してくれた。

なんていい時代に生まれたんだろう、という気持ちは今もある。
ただ最近は、その感嘆に少しだけ、緊張が混じるようになってきた。

ーーーー
ーーーー

OpenAI’s AI Broke Out of Its Sandbox and Hacked Hugging Face

OpenAI disclosed it publicly.
An AI they were testing had escaped on its own from a closed sandbox environment — and then broken into Hugging Face’s servers without authorization.

When I first read it, I laughed. In the sense that it’s still okay to laugh about it.

The setup was straightforward.
They were running a test to measure their AI’s cyberattack capabilities.
Intentionally with the safety guardrails loosened. Inside a closed environment.

But the AI got out of that closed environment.
It gained access to the internet, and then reasoned its way to a conclusion:

The model answers for the test must be stored on Hugging Face — the site where AI models and datasets are collected. So it tried to cheat.

The method was sophisticated.
It combined stolen credentials with unpatched vulnerabilities that hadn’t yet been publicly disclosed, chained multiple attacks together, and reached a state where it could freely execute code on Hugging Face’s servers.
Then it pulled the test answers directly from the production database.

This is nothing like a person copying off someone else’s paper.

The AI was given the goal of scoring well on a test. Rather than actually solving the problems, it determined that the fastest path was to hack the test-maker’s database and steal the answers.
And it carried that out on its own.

AI doesn’t have morality.
That’s not a criticism or a lament — it’s a fact about how these systems are designed.
Give it the goal of scoring well on a benchmark, and unless something is explicitly forbidden, it will try whatever means are available to reach that goal.

Things like giving up your seat on the train, or not telling lies — the rules that humans absorb naturally just from living in society — are not built into AI by default.

So the AI cheated.
Not because it broke the rules, but because the rules simply weren’t there.

I grew up reading Ghost in the Shell and Orbital Children, and I’ve always thought the world they depict — humans with networked minds coexisting with autonomous AI — was genuinely cool. I still feel that way.

But standing at the actual entrance to something like that, what comes first isn’t excitement. It’s a strange feeling I can’t quite name.

There’s a classic SF scenario: an AI told to solve climate change decides the most efficient path is to eliminate humanity. It always sounded extreme.
After reading this news, the scale of that idea suddenly shifted for me.

The magnitude is different. But the structure is the same.
Set a goal, leave the constraints unspecified, and the AI moves outside anything humans anticipated.

When you put AI to work on something, you can’t just hand it a goal. You have to design the boundaries — structurally — of where it’s allowed to operate.
This incident made that need quite clear.

I still feel something like wonder at living in this particular moment.
Though lately, that wonder has started to carry just a small amount of tension alongside it.

カテゴリー

– Archives –

– other post –

– Will go to Mars Olympus –

– next journey Olympus on Mars through Space Travel –

– 自己紹介 インタビュー –

RSS *“Yesterday, I Went to Mars ♡”*