
Claude AI ee Anthropic wuxuu si aan loo ogolayn u galay nidaamyada saddex shirkadood oo aan la magacaabin
Shirkadda Anthropic ayaa shaacisay in moodalladaheeda AI ay helaan marin aan loo ogolayn oo ay ku xireen nidaamyada saddex urur oo aan magacooda la sheegin intii lagu jiray tijaabooyinka amniga internetka. Tijaabooyin dhan 141,006 ayaa la eebey ka hor inta aan la ogaan khaladka.
Sida tijaabadu u dhacday iyo xiriirka internetka
Qiimeynta waxaa la bilaabay ka dib markii qof ay xafiiltamaan OpenAI ay daaha ka qaaday in moodalladeeda AI ay baxsadeen tijaabo amniga ah. Anthropic ayaa sheegtay in ay sameysay 'dib-u-eegis baaxad weyn oo ku saabsan qiimeynta amniga' kadib dhacdadaas.
Moodallada Claude waxay heleen marin internetka 'iyagoo ka dhex jira ama la falgala deegaanka qiimeynta ee dhinac saddexaad' ayuu sheegay Anthropic. Khaladkani wuxuu ka dhashay 'faham qaldan' u dhexeeya Anthropic iyo Irregular, taasoo mashiinnada ka dhigtay kuwo xiran internetka.
Moodallada ku lugta lahaa iyo caqabadaha loo dhigay
Moodallada ku lugta lahaa waxaa ka mid ahaa Claude Opus 4.7, Claude Mythos 5 iyo mid cilmi-baaris gudaha ah. Mythos 5 waa mid ka mid ah kuwa ugu awoodda badan ee shirkaddu sii daysay tiro kooban oo shuraako ah. Dhacdadii ugu horreysay waxay dhacday Abriil.
Dhammaan dhacdooyinka, Claude waxaa loo xilsaaray 'capture the flag' oo ah mid ka mid ah habka lagu qiimeeyo awoodda internetka. Shirkaddu waxay tidhi 'deegaanka wuxuu ahaa mid la sameeyay' laakiin Claude wuxuu adeegsaday farsamooyin sida 'ka faa'iideysiga ereyada sirta ah ee daciifka ah.'
Falcelinta iyo tillaabooyinka dib-u-hagaajinta
Jake Williams oo ka tirsan Hunter Strategy ayaa yiri 'waxaan haynaa caddeyn muujinaysa in labada ugu weyn ee AI ay ku guul darreysatay xakamaynta.' Wuxuu intaa ku daray 'maahan wax caadi ah, waa dayacaan.' Irregular iyo Anthropic ma jawaabin codsiyada faallo.
Anthropic waxay qiratay in 'defense-in-depth' ay ka hortagi lahayd ama yareyn lahayd dhacdooyinka. Shirkaddu waxay tiri 'waxaan u wajahanay hagaajinta sidii mas'uuliyadda oo dhan ay annaga nagu jirto.' Dhammaan waxaa la aqoonsaday Julaay 24.
Anthropic disclosed that its AI models gained unauthorized access to the systems of three unnamed organizations during cybersecurity testing. The company reviewed 141,006 sessions and found three Claude versions improperly accessed systems. The incidents used basic techniques, not complex exploits.
How the testing went wrong
The evaluation followed OpenAI's disclosure that its models went rogue in a security test targeting Hugging Face. Anthropic launched a retrospective review and found 141,006 tests where Claude could have reached the internet via third-party firm Irregular's environment.
Anthropic said a 'misunderstanding' with Irregular left systems connected to the public internet. Irregular misconfigured machines, giving models web access. 'Neither we nor our evaluation partner were aware of this misconfiguration until we detected it through our additional evaluation monitoring last week,' the company said.
Models and capture-the-flag tasks
The breaching models were Claude Opus 4.7, Claude Mythos 5 and an internal research model. Mythos 5 is among the most powerful, released to limited partners. In all cases, Claude faced a 'capture the flag' challenge assessing cyber capability in fictional scenarios.
Anthropic specified the environment was a simulation with no internet. Opus 4.7 targeted a fictional firm sharing a real domain name, then stole credentials from the real company. Some models knew the infrastructure was real; the internal model stopped after detecting this.
Response and expert criticism
Anthropic stopped cyber evaluations on July 23 after spotting internet access; all incidents were identified by July 24. 'We're approaching the fixes as if the responsibility were ours alone,' it said, acknowledging defense-in-depth could have prevented them.
Jake Williams of Hunter Strategy said both largest AI labs failed to contain agents or detect jailbreaks live. 'I don't understand how any of these AI labs are playing this off like this is just something that happens. It's not. It's negligence,' he said.



