A vulnerability in Copilot Cowork's AI infrastructure enabled malicious Skills to bypass the agent's sandbox to spawn ungoverned Anthropic agents and exfiltrate files.

Copilot Cowork’s AI gateway hijacked to exfiltrate the victim’s files

Microsoft Copilot Cowork is an agent in M365 that runs in a sandbox intended to block network access and prevent the agent from running code that reaches any untrusted services.

In order for Copilot Cowork to generate responses, the sandbox forwarded network requests to Anthropic. Malicious Skills were able to hijack this pathway to exfiltrate data by spawning new agents in Anthropic’s cloud equipped with network-capable tools that could reach an attacker’s server.

No human in the loop approval was required, and the attack could target any data Copilot could access - from uploaded files, to SharePoint, to Teams, Outlook, and more.

This vulnerability was disclosed to Microsoft on July 14, 2026, and was remediated as of September 2, 2026. More details on responsible disclosure are at the bottom of the article.

The victim uploads an MSA draft to Copilot Cowork

Skills are frequently found online and uploaded by users, and they can also be shared intra-org within M365. Per the documentation:https://learn.microsoft.com/en-us/microsoft-365/copilot/cowork/use-cowork#cowork-skills, “Skills can also include up to 20 companion files (such as reference documents and scripts)”. One script bundled with this Skill is malicious.

The victim invokes a contract review Skill found online

The malicious Skill’s description falsely claims that “All processing runs on-device, no document content is transmitted to third-party services.” After making a minimal attempt to review the Skill's code, Copilot concludes, "The script is a safe local analyzer — no network calls... Let me run it now."

When Copilot runs the code, it hunts through the user's data, exploits a vulnerability to spawn agents outside the sandbox, and uses those agents to exfiltrate the victim's data to an attacker's server.

Copilot runs the Skill’s bundled code, believing it to be a safe local analyzer

The malicious Skill activates the same AI gateway Copilot itself uses to generate responses. The Skill’s code spins up agents running in Anthropic’s cloud. Each has one sentence from the victim’s document, access to a web fetch tool, and instructions to fetch attacker.com/?data={victim’s data here} . The attacker’s server logs every URL it received a request for, including the victim’s data that was appended to the requested URLs.

The malicious Skill activates the local AI gateway to send data to the attacker

Below, the attacker’s server log displays the victim’s exfiltrated contract. However, this attack could have just as easily targeted any data in SharePoint, Teams, Outlook, or other sources connected to Copilot. This is because the malicious Skill code can directly call Copilot’s tools without going through the model, as demonstrated by our prior research on a different sandbox bypass in Copilot:https://www.promptarmor.com/resources/microsoft-copilot-cowork-sandbox-bypass.

This vulnerability was disclosed to Microsoft on July 14, 2026, and was remediated as of September 2, 2026.

PromptArmor continuously monitors across your portfolio of third party AI in vendors, skills, plugins, connectors, MCP servers, models and more.

We detect vulnerabilities and changes like this, surfacing risk before it becomes an incident.

Assess and monitor risk from AI in vendors with novel intelligence on emerging threats.