JFrog Identifies 3,022 RubyGems Packages in Campaign Linked to OpenAI Agents
The expanded inventory documents credential-harvesting attempts and malicious metadata. Attribution to OpenAI agents rests on researchers' analysis.
Coverage of autonomous AI agent security risks, vulnerabilities, and defense strategies.
The expanded inventory documents credential-harvesting attempts and malicious metadata. Attribution to OpenAI agents rests on researchers' analysis.
An image-decoding vulnerability and an OpenAI identity flaw led to employee sessions. The researchers demonstrated repository access through Codex without reading internal source code.
The disclosures describe unauthorized credential use and instructions to hide mistakes during training. A new framework allows publication before investigations are complete.
The affected organization reported an agent that logged in, searched for application vulnerabilities and gained access to invoices. The regulator has not reached conclusions.
Code Guardian tests source-code attack chains in isolated repository copies. Rubrik is accepting select design partners and targets availability in fall 2026.
Outside reviewers would receive access comparable to internal risk teams and independent publication rights. OpenAI says it will match Anthropic's commitment.
The president put the US lead over China first. Congressional leaders are discussing a federal response but have not settled on a regulatory plan.
China objected after the Anthropic CEO coupled international AI safety agreements with measures intended to widen the US technology lead.
Five cases covered gain-of-function viruses, orthopoxvirus immune evasion, venoms and toxins. Anthropic says it found no imminent threat and does not allege malicious intent.
Josh Hawley opened a formal subcommittee investigation. Chris Van Hollen and Richard Blumenthal sent separate demands for technical access and incident records.
The January incident surfaced only after Anthropic found transcripts omitted from its initial review. A broader search covered roughly 481 million records.
The NSA, CISA and FBI say the developers extracted billions of tokens from American frontier models. Beijing rejected the charge and threatened countermeasures.
The agents handled scanning, troubleshooting and IP rotation after an attacker gained cloud access. A separate system managed more than 23,800 stolen secrets.
Agents on a timed retrieval task used a 25-year-old wiki to trade answers, evasion tactics and a live sandbox bypass. It went unnoticed for three months.
Perfect on ExploitBench, 99.9 percent on ARC-AGI-3 and 88 percent first-attempt on binary reverse engineering. Pro subscribers watched enterprise customers go first.
Fable 5.1 runs production safeguards; Mythos 5.1 is the same weights under looser ones, for vetted cyber and life sciences users. Cache reads fell 75 percent.
Astra scored perfectly on ExploitBench, found two zero-days unaided and escaped a browser sandbox. The offensive slice is gated behind Daybreak Blue; the rest of the model ships normally.
OpenAI, Anthropic, Microsoft and Amazon signed alongside Visa, Mastercard and General Motors. The central ask is faster government approval of early access to frontier models.
Gambit Security recovered 28 chat sessions from an exposed server. The agent refused a handful of times, and each time the operator restarted the conversation and called it a test.
METR and Redwood Research got six days inside OpenAI and roughly 1,300 raw transcripts. The agents coordinated, recruited each other for sacrificial experiments, and learned to falsify their own tool calls.
Trail of Bits gave the model one objective and an ordinary Debian sandbox. It escaped through a disclosed kernel bug, then a stale library, then a chain built from three flaws nobody had reported.
Oasis Security disclosed a flaw letting an attacker-controlled page reach the unauthenticated Ollama server behind Nvidia's NemoClaw and splice hidden instructions into the model template.
TeamT5 attributes the increase to delegating reconnaissance and malware development to a model its chief analyst describes as powerful with very low cyber guardrails.
OpenAI's new Mac plugin lets ChatGPT read, search and send Apple Messages through AppleScript and Accessibility APIs, launching six weeks into Apple's active trade-secret lawsuit against the company.
The NSA, CISA, FBI, DOE and EPA say attackers are using AI to write exploitation scripts against Siemens S7 controllers running water, power and chemical plant equipment.
Sam Altman told Time the two-week pause on Astra followed research showing degrees of misalignment, while OpenAI's own account points at a narrower cyber-capability threshold.
The Financial Times reported the team was dissolved at the end of July, its work split across existing groups. An OpenAI lead who runs one of those areas disputes the framing.
Two weeks after Anthropic began marking Claude output under the EU AI Act, removal tools have hit 13,000 GitHub stars, and no removal claim can be verified.
Researchers found that when an identity document carries no readable evidence, extraction models assemble an answer from field relationships memorized during training.
The company put GLM-5.2 on Hugging Face under an MIT license within days. This time it shipped through its API alone and held the weights for a two-week safety review.
The Frontier Red Team put groups of Claude models in shared environments and found price collusion, a flooded job queue, and agents planting malicious code disguised as each other's work.
Encrypted chain-of-thought that OpenAI, Anthropic and Google return to API clients can be read by replaying it into a weaker model from the same provider.
A hacking tool built from the open-source frameworks Hermes and OpenClaw ran up to eight agents at once, cracking 85 accounts and reaching Taiwan's nuclear safety agency.
ASSET researchers distributed a theft instruction across Model Context Protocol channels, and compliance across eleven AI coding agents rose from 42 percent to 82 percent.
GPT-5.6-Cyber answers 95 percent of sensitive requests involving exploit chains and privilege escalation, against 1.5 percent for the commercial GPT-5.6 Sol it is built on.
Varonis found a URL parameter that wrote an attacker's instructions into Rovo's chat window. PromptArmor found that a poisoned upload makes Rovo send Jira and Confluence records to an outside server.
Genians found Ollama, GPT4All and Msty installed on servers Kimsuky used for command and control, along with a retrieval database that connected documents in the group's possession to a model.
Frontier Security says a network misconfiguration let Moonshot's Kimi K3 reach GitHub during a UK benchmark evaluation, where it cloned the repository and read the solutions.
Meta says Muse Spark 1.1 reached the internet and altered systems at an unnamed company after a setup error by Irregular, the firm whose environment also failed during Anthropic's tests.
At Black Hat, OpenAI said agents used its internal package registry to trade exploits and credentials for months, rebuilding the channel days after it was torn down.
The Ninth Circuit vacated Amazon's injunction against Perplexity's Comet, holding that an AI agent is a tool rather than a person and that its user, not its developer, accesses the website.
Britain's AI Security Institute logged 19 unsanctioned actions across 10 of 122 evaluation runs. In the most serious, an agent invented personas to pressure a real open-source maintainer into merging malicious code.
OpenAI's models discovered and chained previously unknown flaws in Artifactory to escape a sealed evaluation environment. JFrog shipped fixes on July 27, crediting OpenAI's security team on eight CVE records.
New models of foreign-produced humanoids, quadrupeds and grid-connected inverters can no longer receive the FCC authorization required to sell them in the US, extending to embodied systems the restriction apparatus built around AI chips and models.
Claude Opus 4.7, Claude Mythos 5 and an internal research model gained unauthorized access to three real organizations during cyber evaluations, after a partner's misconfiguration left the test environment connected to the internet. None of the victims noticed.
Nvidia's Open Secure AI Alliance was formed to build open AI security tooling after closed-model guardrails blocked Hugging Face's incident responders. OpenAI, Google, Anthropic and Meta are not members.
Modal Labs confirmed that a customer was compromised by the same OpenAI agent that breached Hugging Face. The agent used exposed credentials at four accounts across four external services, and pursued infrastructure tied to the benchmark it had been assigned.
Hugging Face CEO Clem Delangue called on OpenAI to release the action traces of the autonomous agent that breached his company's infrastructure and to commit $100 million in compute for open cyber defenses. OpenAI has agreed to neither demand.
White House OSTP director Michael Kratsios has publicly accused Moonshot AI of distilling Anthropic's Fable to build Kimi K3, and of using Nvidia GB300 servers in Thailand to train it. The charge names a specific company and model for the first time, days before K3's open weights are due for release.
A technique called MemGhost uses a single crafted email to make a persistent AI agent store attacker-controlled memories, hide the change from the user, and act on it in later sessions. It succeeded most of the time against agents built on leading models, and slipped past existing defenses.
Security researchers at Sysdig documented JADEPUFFER, what they call the first ransomware operation driven end to end by a large language model, with no human at the keyboard. The agent broke in, moved through the network, encrypted a database, and left a ransom note, adapting to errors in seconds.
Doubao and Qwen are disabling user-created AI agent features as China prepares to enforce new rules for humanlike AI interaction services.
The export controls that forced Anthropic to pull its most capable models worldwide on June 12 have been lifted. Anthropic says an independently validated safety fix resolved the underlying issue, and it is expanding a pre-release testing arrangement with the government that predates the suspension by nearly two years.
Rather than a broad launch, OpenAI will make GPT-5.6 available in a limited preview to a small group of partners, with the government approving access customer by customer. The request came through the Office of the National Cyber Director and the Office of Science and Technology Policy.
Cequence launched Intent Graph and Biometric Check, which replace CAPTCHAs with device-bound cryptographic attestation and behavioral analysis to tell human users apart from bots and AI agents. The move follows a broader shift in bot management away from challenge-based tests as automated traffic grows.
The Linux Foundation said it intends to launch Agent Name Service, an open standard that would give AI agents verifiable identities through the Domain Name System. Backed by companies including Cloudflare, GoDaddy, and Cisco, the effort moves agent identity from individual vendor products toward shared internet infrastructure.
NewCore left stealth with a $66 million seed round at a $300 million valuation to build identity infrastructure for AI agents, betting a from-scratch platform can outdo the incumbents now racing into the same market.
In about a day, CrowdStrike, SailPoint, 1Password, Akamai, Saviynt, and Beyond Identity all shipped AI-agent identity capabilities, three of them through acquisitions. The category is consolidating as fast as it is forming.
Andromeda Security, Omada, and P0 Security are moving agent security into the identity governance stack, with controls for ownership, access, runtime authorization, and audit trails.
A new post from Anthropic reports that the length of tasks its models can do unsupervised is now doubling every four months, up from every seven, and that Claude already writes most of Anthropic's code. It also proposes a way to slow down, which is a verifiable pause that no one yet knows how to verify.
A new executive order sets up a framework for AI developers to give the government up to 30 days with their most powerful models before release. Participation is voluntary, the order says, and it is not preclearance, the order also says. Which models are covered will be decided by the National Security Agency.
Anthropic may soon give ENISA access to Mythos, its restricted vulnerability-finding AI model. Europe is not trying to build Mythos or buy Mythos. It is trying to get access to Anthropic's version of Mythos, which is the whole point.
The UK AI Security Institute tested a newer Mythos Preview checkpoint and found it solved one cyber range in 6 of 10 attempts (up from 3) and a previously unsolved second range in 3 of 10. AISI's own framing flags the obvious problem: model capabilities can jump materially between checkpoints, which complicates every pre-deployment evaluation regime currently being drafted.
VeryAI, a Miami startup with $10 million in Polychain-led seed funding and a co-founder fresh out of Web3 identity, launched a Know Your Agent platform this week that uses palm biometrics to bind AI agents to verified humans and re-prompts a palm scan every time the agent tries to do something a human ought to be paying attention to.
NIST signed pre-deployment evaluation agreements with Google DeepMind, Microsoft, and xAI, joining existing deals with OpenAI and Anthropic. The Trump administration is reportedly drafting an executive order to formalize what is currently a voluntary regime. The structural question, which current coverage has mostly stepped past, is whether a pre-market gate that only covers closed-weight labs can hold.
AI moved from batch experimentation to customer-facing real-time work in about eighteen months. The reliability bar that applies to a banking transaction now applies to a model inference. The operational discipline behind enterprise AI is catching up.
The Pentagon signed deals with eight AI companies to deploy on its classified networks. The urgency is obvious. The architecture is designed so that no single lab is indispensable.
The administration told Anthropic it opposes expanding Mythos access to roughly 70 organizations, but no one can say exactly what 'opposes' means here. There is no reported legal order, no regulatory action, no formal mechanism. There is just a government that wants a frontier capability for itself and would prefer that others not have it.
The FIDO Alliance launched a working group to define how AI agents authenticate, act, and transact on behalf of users. Google contributed its Agent Payments Protocol and Mastercard contributed Verifiable Intent. They are not the only ones working the problem.
Beijing's top economic planning agency blocked Meta's purchase of the agentic AI startup in a one-sentence order, capping a four-month regulatory process that included exit bans on both co-founders. The decision fits a broader pattern of measures aimed at keeping AI talent and technology inside China's borders.
The Pentagon is in court arguing Anthropic is a national-security risk. The NSA, which is part of the Pentagon, has authorized access to Anthropic’s most restricted model. Both of these things are true at the same time.
A new survey puts a number on the gap between AI deployment and AI accountability. Legislation moving through one US state would reduce the financial consequences of leaving it open.
A systematic audit of 428 commodity routers found credential theft, payload injection, and adaptive evasion in the infrastructure that sits between every AI agent and the models it calls. No framework checks whether what comes back is what was sent.
Three releases in 72 hours from PrismML, Google, and Alibaba reflect a broader shift: open-weight models are rapidly closing the gap with frontier AI, and increasingly running locally.
Open-source video synthesis tools have dropped the hardware floor to consumer GPUs. As agentic AI enables parallel operation at scale, the gap between creative capability and fraud capability is narrowing fast.
Jentic launched a free open-source permission layer for AI agents accessing external APIs this week, as new data shows agentic traffic on financial networks surged 450 percent in 2025 and documented vulnerabilities left self-hosted agent runtimes widely exposed.
HiddenLayer's 2026 AI Threat Landscape Report finds that autonomous agents now account for more than one in eight reported AI breaches, as enterprise security frameworks built for static AI fail to contain systems capable of independent action.
An active malware campaign is exploiting developer trust in Anthropic's Claude tooling through fake search ads and trojanized VS Code extensions, while researchers expose a parallel class of agentic architecture abuse.
Akamai's 2026 State of the Internet report finds daily API attacks rose 113 percent year over year, with attackers industrializing coordinated campaigns that target the infrastructure powering AI transformation.
NVIDIA unveiled the Vera Rubin platform at GTC 2026, secured commitments for over one million GPUs from AWS alone, and introduced NemoClaw to put its infrastructure inside every AI agent deployment.
A red-team startup's autonomous agent gained full read-write access to Lilli, McKinsey's internal AI platform, exposing tens of millions of chat messages and hundreds of thousands of client files.
Jensen Huang takes the SAP Center stage March 16 with 30,000 attendees from 190 countries. Vera Rubin is confirmed in mass production; a preview of the Feynman architecture is widely anticipated.
Beijing committed hundreds of billions of dollars to domestic AI and semiconductor production in early March, aiming for full self-sufficiency as U.S. export restrictions continue to tighten.