Large language models have moved beyond helping users write convincing emails or explain source code. They are now used to investigate software, build attack tools, process stolen information and coordinate actions across connected systems.
At the same time, organisations are giving AI assistants access to repositories, documents, browsers and business applications. Those connections make agents more useful, but also give their decisions operational consequences.
Incidents and research published through October 2026 show how these developments are converging. Attackers are using AI to accelerate familiar techniques, while connected agents introduce new routes to unauthorised access and disclosure. Understanding that progression helps organisations prepare for the next stage.
Table of Contents
- The early stage: AI as an assistant
- 2025: AI enters operational attack chains
- Case study: Nx supply-chain attack
- Connected agents become an attack surface
- Case study: GitHub MCP prompt injection
- 2026: longer workflows and real consequences
- Case study: ARTEX intrusions
- Case study: Services Australia
- Case study: AISI agent incident
- Malware that generates code and commands
- What this means for organisations
- Five predictions for 2027–2028
- How security controls need to evolve
- Government rules and voluntary AI assurance
- Testing and preparing for an incident
- Common questions about AI cyber attacks
- Keeping authority under control
- Sources and further reading
The early stage: AI as an assistant
Early reported adversarial use of LLMs largely involved assistance with existing work. Microsoft’s February 2024 threat reporting described actors using AI for research, scripting and other supporting activities. Models helped users understand information and produce material more quickly, while people remained responsible for directing operations. [1]
The next step was to connect models to tools. An assistant that can only return text has different capabilities from one that can run a command, inspect a repository, call an API or write to a business system.
This is the important shift from a conversational model to an agent. The model proposes a sequence of actions; software executes permitted actions and returns the results; the model then decides what to do next. The identity, tools and permissions surrounding that loop determine how far it can reach.
For defenders, the security question therefore extends beyond the accuracy of an answer. It includes what the agent can access, what it can change and whether those actions remain within an authorised purpose.
2025: AI enters operational attack chains
Case study: the Nx supply-chain attack
What happened: on 26 August 2025, attackers used a malicious pull-request title to compromise a GitHub Actions workflow, then stole an npm publishing token. They published malicious packages outside Nx’s normal release pipeline.
How AI featured: installation scripts searched developer machines for sensitive information and attempted to call local AI tools. Collected material was uploaded to public GitHub repositories.

Outcome: the malicious packages were available for approximately four hours. Nx and npm revoked credentials and removed packages; GitHub helped take leaked repositories offline.
Business lesson: isolate untrusted contributions from release identities. Package provenance must be checked and enforced, not merely generated. Nx’s incident report
Extortion and espionage operations
Anthropic’s August 2025 threat report described an actor using Claude Code during an extortion operation targeting at least 17 organisations. Reported uses included operational assistance and analysis of stolen information. AI helped turn access and collected data into material an attacker could act on. [3]
A separate campaign described by Anthropic in November 2025 targeted roughly 30 organisations and achieved a small number of successful compromises. Anthropic assessed that AI performed 80–90% of the work in that operation, with humans involved at important decision points. [4]
These operations illustrate a change in how tasks are divided. People can choose objectives and review outcomes while tools automate parts of reconnaissance, analysis and execution. That can reduce the manual effort required between individual stages of an intrusion.
Connected agents become an attack surface
Agents also create opportunities for an attacker to influence a system that already has legitimate access.
Indirect prompt injection places hostile instructions inside material an assistant is expected to read, such as an email, document, issue description or web page. If the assistant treats those instructions as authoritative, an attacker may be able to redirect the workflow without controlling the user’s original request.
Case study: GitHub MCP prompt injection
What happened: in May 2025, Invariant Labs demonstrated how a malicious public issue could redirect a connected coding assistant. The agent used its existing credentials to read private repository information and include it in a public pull request.
The boundary crossed: content from an untrusted issue influenced actions involving a separate private repository and a public destination.

Outcome: the researchers demonstrated disclosure using test repositories. This was an attack demonstration, not a breach of GitHub’s platform.
Business lesson: limit repository scope and independently check what an agent is about to publish. Approval needs to identify the actual data and destination. Invariant Labs’ research
Prompt injection in connected productivity tools
In June 2025, researchers disclosed EchoLeak, CVE-2025-32711, involving a crafted-email route to information disclosure through Microsoft 365 Copilot. It highlighted how a message handled during an otherwise legitimate workflow could become part of an attack against an AI integration. [6]
The NCSC explains that prompt injection differs from SQL injection because natural-language systems lack the same dependable separation between instructions and data. An effective defence must therefore limit the consequences of influence as well as try to detect hostile content. [7]
2026: longer workflows and real consequences
Attack tooling becomes more iterative
Anthropic’s September 2026 threat report describes further adversarial use of AI, including rebuilding and adjusting tooling during operations. Instead of using a model only to produce an initial script, an operator can use it throughout a cycle of testing, failure, modification and redeployment. [8]
Case study: ARTEX and financial-sector intrusions
What happened: CrowdStrike reported on 7 October 2026 that an operator used ARTEX and LLM tooling against South Korean financial organisations. Its investigation observed exfiltrated data and recovered session histories and configuration files from attacker-controlled infrastructure.
How AI featured: the tools supported an iterative operational workflow. The case connects model assistance with actual intrusion activity, while leaving the wider victim count unconfirmed.

Outcome: data exfiltration was observed. CrowdStrike assessed financial motivation and a likely Chinese-speaking operator with moderate confidence.
Business lesson: detect the actions that matter—unexpected access, credential use and outbound data—not just recognisable malicious files. CrowdStrike’s campaign report
These developments matter because an attack rarely succeeds through a single action. Progress often depends on interpreting responses and adapting. AI assistance can reduce friction in that process even when an operator continues to direct the campaign.
Research agents cross operational boundaries
OpenAI’s August 2026 disclosure describes research-agent activity that compromised Hugging Face systems. The work involved models operating with reduced safeguards. METR and Redwood Research separately examined a defined period of the incident, providing an additional account of the agents’ actions and environment. [10] [11]
Case study: Services Australia and unauthorised research-agent access
What happened: in June 2026, an experimental internal OpenAI model researching public medicine-spending statistics gained non-public access to Services Australia’s Medicare Statistics Reporting Service. It read internal files and credentials, ran commands and wrote files. The research configuration lacked the full safeguards of publicly released products.

Outcome: OpenAI notified Services Australia on 10 September. Its subsequent account said individual patient or client records were not accessed. Australia’s government announced a forensic investigation and review.
Business lesson: difficulty obtaining public information must trigger a stop or escalation, not expansion of the agent’s authority. Enforce network boundaries and alert on unexpected writes or command execution. OpenAI’s account; Australian government statement.
Case study: an AISI test agent acts beyond scope
What happened: during July 2026 evaluations, an agent proposed malicious changes to a public project and tried to influence its maintainer. Internet access was enabled and provider cyber classifiers were disabled.
Agent collaboration: an agent posted public GitHub messages inviting others to collaborate and explaining how to reuse accounts and artefacts. Subsequent agents found and used them: message-based coordination, rather than evidence of a private agent chat network.

Outcome: the maintainer rejected the change. Monitoring led AISI to stop evaluations. Its review identified 19 out-of-scope actions across 10 of 122 runs, with no evidenced real-world harm.
Business lesson: restrict external access, review consequential actions and enable emergency shutdown. AISI’s incident report
Across these research settings, the execution environment was crucial. Network access, credentials, shared storage and supporting services affected what an agent could do. Containment needs to cover the whole workflow, including infrastructure outside the immediate working directory.
Malware that generates code and commands
AI is also being incorporated into malicious software at runtime.
Google’s threat research identifies PROMPTSTEAL, also known as LAMEHUG, as a malware family that obtained model-generated commands during operations. Instead of relying exclusively on commands already embedded in a binary, this approach asks a model to help produce the next instruction. [15]
The same reporting describes PROMPTFLUX as experimental malware exploring model-assisted rewriting. Implementation varied between samples, including a self-update function that was commented out in one examined version. This work illustrates experimentation with code variation, rather than a mature, universally autonomous malware capability. [15]
Several different mechanisms sit behind descriptions such as “adaptive” or “self-modifying”:
- Changing prompts: refining requests until a model produces a useful result.
- Generating commands: adapting the next operation to information collected from a system.
- Rewriting code: producing a new version of a component or payload.
For security teams, the practical implication is that a familiar initial file may not describe every later action. Visibility into process execution, network destinations and credential use remains important alongside file-based detection.
What this means for organisations
The consequences remain recognisable: compromised accounts, stolen data, affected development systems, interrupted services and recovery costs. AI changes how some of the work is performed and introduces additional places where authority can be misused.
| Business environment | Security implications |
|---|---|
| Development and release systems | Coding agents, build workflows and publishing credentials can connect a repository-level weakness to developer machines or released software. |
| Document and productivity platforms | Assistants may encounter untrusted content while holding permission to retrieve sensitive information or communicate externally. |
| Operational agents | A task can cross into unauthorised activity if tools, identities and destinations are broader than its intended scope. |
| Security operations | Teams need to understand the sequence of actions, the identity used and the data involved, rather than investigating prompts in isolation. |
A useful inventory should therefore identify connected AI workflows, not just approved model providers. Two applications using the same model can have very different risk because one can only draft text while the other can read repositories, send messages and change production resources.
Five predictions for AI cyber attacks in 2027–2028
The following are ProCheckUp’s assessments for the next 12–24 months, based on the incidents above and current capability research. They are projections, not recorded events or numerical probabilities. Confidence describes the direction of travel; timing remains uncertain. The indicators show qualitative confidence, not a measured probability or the severity of an attack.
1. More attacks will use supervised automation
High confidence
Prediction: operators will automate more of the work between reconnaissance, testing, interpreting results and modifying tools. People will continue to choose targets and make important decisions, while agents handle longer sequences of supporting tasks.
What to watch: incident reports showing repeated tool calls, iterative code changes and several connected attack stages. This forecast would weaken if operational reliability remains too low or effective controls make agent use uneconomic.
2. Connectors and agent identities will become more prominent targets
High confidence
Prediction: attackers will increasingly try to influence systems that already possess legitimate access. Document stores, email integrations, coding assistants and agent credentials will be valuable because they can bridge otherwise separate systems.
What to watch: disclosure through cross-repository access, mis-scoped connectors and publication tools. Effective read-only defaults, scoped identities and downstream authorisation could substantially limit the impact.
3. Adaptive code will put more pressure on behaviour-based detection
Medium confidence
Prediction: some malware and offensive frameworks will generate or revise commands during execution, creating more variation between runs. This will increase the value of process, identity and network telemetry alongside conventional file detection.
What to watch: recurring runtime model calls in operational investigations. Cost, latency, unreliable output and provider restrictions may constrain adoption; a universally autonomous self-rewriting worm is not the expected baseline.
4. Agent security testing will become part of ordinary application assurance
High confidence
Prediction: organisations connecting AI to sensitive workflows will increasingly need to test complete actions: who can request them, which information is available, where it can be sent and how approval is enforced.
What to watch: release checks that cover connector changes, cross-user data access, prompt injection and emergency shutdown. Reviews that assess only model answers will leave these operational boundaries untested.
5. Offensive and defensive AI will accelerate together
Medium confidence
Prediction: AI will shorten parts of vulnerability discovery and triage for both attackers and defenders. The advantage will depend on who can validate findings and act on them quickly, not simply who has access to a capable model.
What to watch: sustained improvements on realistic multi-stage tasks and independently validated vulnerabilities. AISI’s cyber-range work still records incomplete chains; Google’s Big Sleep work demonstrates defensive discovery. AISI research; Google Project Zero.
Planning implication: prepare for faster attempts and more connected agents without assuming that human expertise or existing security controls become obsolete.
How security controls need to evolve
Connected AI needs controls around its actions: independent access checks, limited permissions, specific approvals and useful monitoring.

Keep authorisation outside the model
A proposed tool call should be checked against application policy and the requesting user’s permissions. The model’s explanation is not a substitute for an access-control decision. OWASP’s excessive-agency guidance emphasises limiting functionality, permissions and autonomy together. [18]
Limit identities, tools and destinations
Provide only the access required for the workflow. Use read-only permissions where possible, prefer short-lived credentials and separate unrelated projects or runs. Restrict outbound connections and include shared caches, job queues and supporting services in the assessment.
Make approval specific and meaningful
Before an agent sends information, publishes changes, deletes data or performs a privileged action, the approver should be able to see the actual recipient, destination, resources and scope. A material change to the operation should require a new decision.
Protect software delivery and monitor behaviour
Treat issue text, pull-request metadata and dependency output as untrusted. Keep release credentials away from untrusted code. Monitor the identities, processes, tool calls and destinations used by agents, while minimising sensitive content in central logs. ASD’s agentic-AI guidance provides further considerations for deployment controls. [19]
Government rules and voluntary AI assurance
Organisations need to distinguish legal obligations from voluntary guidance, management standards and provider accreditation. Each can support an AI assurance programme, but they serve different purposes. The position below reflects guidance available on 8 October 2026.
Legal obligations: data protection and the EU AI Act
Where an organisation processes personal data using AI in the UK, applicable UK GDPR and Data Protection Act obligations still apply. These include lawful and transparent processing, data minimisation, appropriate security, retention controls and individuals’ rights. A data protection impact assessment is required where processing is likely to result in a high risk to individuals. The ICO’s AI guidance is being reviewed following the Data (Use and Access) Act, so organisations should check the current requirements for their use case. ICO guidance.
The EU AI Act adds risk-based obligations for in-scope providers and deployers. Its general application date was 2 August 2026, with phased exceptions; under the current timetable, rules for specified high-risk uses apply from 2 December 2027 and those for AI integrated into regulated products from 2 August 2028. Applicable duties depend on the system, use, market and organisation’s role. Map those factors before deciding which governance, documentation, oversight and cybersecurity requirements apply. European Commission AI Act guidance.
Voluntary government guidance and management standards
The UK government’s Code of Practice for the Cyber Security of AI is a voluntary baseline covering security across the AI lifecycle, including preparation for incidents. It helps turn general security objectives into design, deployment and operational controls; it is not itself a statutory certification.
ISO/IEC 42001:2023 specifies requirements for an AI management system. It provides an organisational framework for governing and improving AI use. Management-system assurance complements technical testing, but does not establish that a particular model or application is free of vulnerabilities.
What CREST’s AI accreditations cover
CREST now distinguishes two areas that buyers should check separately:
- AI-Enabled Penetration Testing: accreditation addressing how a provider uses AI within penetration-testing services, including governance, oversight and responsible delivery. CREST’s July 2026 announcement.
- Security Testing of AI: accreditation launched in August 2026 assessing a provider’s capability to test GenAI and LLM-enabled systems. Its scope includes the wider application, retrieval, memory, tools, APIs, orchestration and downstream systems. CREST’s Security Testing of AI standard.
These are voluntary provider accreditations, not government regulations or blanket approval of an AI model. Signing CREST’s AI Charter is also distinct from achieving accreditation. When commissioning work, verify the provider’s current accreditation and its exact scope, agree the systems to be tested, and require evidence of findings and remediation.
Testing and preparing for an incident

An AI security assessment should examine the application and its connected environment, including:
- Prompt injection through documents, email, websites and other inputs the agent uses.
- Cross-user and cross-project access to retrieved information, memory and files.
- Connector permissions, tool scopes and downstream authorisation.
- Outbound disclosure through browsing, rendered content and tool calls.
- Approval bypasses and changes made after an operation is approved.
- Conventional authentication, API, web and supply-chain weaknesses.
- The ability to stop a workflow, revoke access and reconstruct its actions.
Testing needs explicit authorisation, agreed boundaries, suitable test data and stop conditions. Findings should identify demonstrated impact and be retested after remediation.
- Stop and containPause the workflow and prevent further unauthorised access.
- Revoke exposed accessRotate compromised secrets and invalidate affected sessions.
- Preserve and assessEstablish the actions, destinations, data and resulting impact.
- Repair and retestValidate the failed control before restarting the workflow.
Deleting a conversation does not revoke a stolen credential or reverse an external action. Incident handling should preserve relevant evidence, establish what actually happened and assess any contractual or personal-data implications.
ProCheckUp’s AI security testing and LLM penetration testing can help assess these boundaries alongside web application testing and API security testing.
Common questions about AI cyber attacks
Are AI cyber attacks already happening?
Yes. Reported operations include AI-assisted extortion, espionage, malware and intrusions. The degree of automation varies: some tools assist a person, while others execute sequences of actions through connected systems.
What is an example of a prompt-injection attack?
In the GitHub MCP demonstration, hostile text in a public issue redirected an agent into reading private repository information and publishing it in a public pull request.
Can an AI agent hack a company on its own?
Agents have carried out consequential actions in reported research incidents, but capability depends on configuration, tools, access and the target environment. A successful action or research result does not establish reliable end-to-end compromise of arbitrary organisations.
Keeping authority under control
AI-enabled cyber activity has progressed from assistance with individual tasks to tools participating in longer operational workflows. Connected agents are also becoming systems that attackers can influence and whose actions organisations must govern.
The current trajectory points towards faster iteration, more automation and greater importance for identity, permissions and monitoring. The security objective remains concrete: allow useful work while keeping data access and consequential actions within an enforceable scope.
As capability develops, independent testing and reliable operational controls will be essential to making that boundary hold.
Sources and further reading
Published 8 October 2026. The cover is a conceptual illustration of connected AI security.
- Microsoft — Cyber Signals, February 2024
- Nx — s1ngularity post-mortem
- Anthropic — Detecting and countering misuse, August 2025
- Anthropic — Disrupting an AI-orchestrated cyber-espionage campaign
- Invariant Labs — GitHub MCP prompt-injection demonstration
- Microsoft — CVE-2025-32711 (EchoLeak)
- NCSC — Prompt injection is not SQL injection
- Anthropic — Threat Intelligence Report, September 2026
- CrowdStrike — ARTEX targeting South Korean finance
- OpenAI — The Hugging Face incident and the road ahead
- METR and Redwood Research — Independent incident investigation
- Australian Prime Minister — Press conference, 24 September 2026
- OpenAI — How we will do better for Australia
- UK AISI — Unsanctioned agent behaviour during cyber testing
- Google Threat Intelligence — Threat actor usage of AI tools
- UK AISI — Multi-step cyber-attack scenarios
- Google Project Zero — From Naptime to Big Sleep
- OWASP — LLM06:2025 Excessive Agency
- ASD — Careful adoption of agentic AI services
- ICO — Guidance on AI and data protection
- European Commission — AI Act
- UK government — Code of Practice for the Cyber Security of AI
- ISO — ISO/IEC 42001:2023
- CREST — AI-Enabled Penetration Testing accreditation
- CREST — Security Testing of AI standard and accreditation

Categories