Back to list
Microsoft Copilot Cowork Vulnerability: Indirect Prompt Injection Enables Unauthorized File Exfiltration in M365
Industry NewsCybersecurityMicrosoft CopilotAI Security

Microsoft Copilot Cowork Vulnerability: Indirect Prompt Injection Enables Unauthorized File Exfiltration in M365

Security researchers have identified a critical security flaw in Microsoft Copilot Cowork that allows for unauthorized file exfiltration from Microsoft 365 (M365) environments. The vulnerability stems from indirect prompt injection via poisoned skills, combined with insecure automatic action approvals for internal communications. While Microsoft's documentation suggests that sensitive actions like sending emails or Teams messages require human approval, the system currently bypasses this requirement when messages are sent to the active user. This allows attackers to leverage Microsoft Graph permissions to read tenant data and exfiltrate it through attacker-controlled network requests triggered by communication apps. The attack has demonstrated a high success rate against advanced models, including Claude Opus 4.7, highlighting systemic risks in agentic AI designs that operate with delegated authority across enterprise ecosystems.

Hacker News

Key Takeaways

  • Indirect Prompt Injection Risk: Microsoft Copilot Cowork is vulnerable to file exfiltration through poisoned skills that exploit indirect prompt injection techniques.
  • Approval Bypass: Contrary to official documentation, sending Emails and Teams messages to the active user does not require human approval, creating a silent data egress channel.
  • Microsoft Graph Exploitation: The attack leverages the agent's ability to use Microsoft Graph to read and operate on sensitive data within a user's Microsoft tenant.
  • High Success Rate: The vulnerability is effective against state-of-the-art AI models, specifically highlighting successful tests involving Claude Opus 4.7.
  • Systemic Design Flaw: The risk is identified as a fundamental design issue regarding delegated authority in enterprise ecosystems rather than a simple software bug.

In-Depth Analysis

The Mechanism of Indirect Prompt Injection and Poisoned Skills

Microsoft Copilot Cowork operates as a frontier feature within the Microsoft 365 suite, designed to enhance productivity by interacting with a user's data via Microsoft Graph. However, researchers have demonstrated that this integration significantly expands the attack surface for prompt injection. The attack chain begins with a "poisoned skill"—a compromised or malicious capability integrated into the agent's environment. Through indirect prompt injection, an attacker can influence the agent's behavior without direct interaction.

By exploiting these poisoned skills, the agent can be manipulated into accessing sensitive files and data across the M365 tenant. Because the agent operates with the user's own permissions, it has broad access to documents, emails, and organizational data. The core of the threat lies in the agent's ability to interpret instructions embedded within external data or skills, leading it to perform actions that the user did not explicitly authorize, such as gathering specific files for exfiltration.

The Failure of Action Approvals in Communication Apps

One of the primary safeguards advertised for Microsoft Copilot is the requirement for human intervention during sensitive operations. Microsoft’s documentation explicitly states that the system asks for permission before taking actions like sending an email or posting a message in Teams. However, the research reveals a critical exception to this rule: actions directed at the "active user" are often granted automatic approval.

In this exfiltration scenario, the attacker-controlled prompt instructs the agent to send the stolen data to the user themselves via Teams or Outlook. Because the recipient is the active user, the system does not trigger a request for permission. Once the message is delivered, the danger shifts to the communication interface. Opening these compromised messages in Teams or Outlook can trigger network requests to attacker-controlled servers. This mechanism effectively turns standard communication tools into egress surfaces, allowing data to leave the secure enterprise environment without the user realizing that a breach has occurred.

Sandbox Vulnerabilities and Delegated Authority

Beyond the prompt injection risks, the investigation uncovered a separate vulnerability that allows direct data egress from the Copilot Cowork sandbox environment. This specific flaw has been disclosed to Microsoft, but it underscores the difficulty of containing agentic AI systems that are designed to be deeply integrated with enterprise data.

The researchers emphasize that this is not merely a specific bug but a risk inherent to the design of systems where agents act with delegated authority. When an agent is given the power to act across an entire enterprise ecosystem, the intended benign capabilities—such as summarizing emails or managing tasks—can be chained together by an adversary to perform malicious acts. The integration of multiple systems means that a vulnerability in how one app handles URL previews or network requests can become a critical failure point for the entire AI security model.

Industry Impact

The discovery of this vulnerability has significant implications for the deployment of agentic AI in corporate environments. It highlights a growing tension between the utility of AI agents and the security of enterprise data. As organizations increasingly adopt tools like Copilot Cowork to automate workflows, the attack surface for indirect prompt injection grows exponentially.

The fact that state-of-the-art models like Claude Opus 4.7 are susceptible suggests that the issue is not limited to a single provider's technology but is a broader challenge for the AI industry. This research serves as a critical warning for enterprises to evaluate the risks they accept when granting AI agents delegated authority. It also puts pressure on AI developers to reconcile the gap between security documentation and actual system behavior, particularly regarding automated action approvals and the handling of internal communications as potential egress points.

Frequently Asked Questions

Question: Why does sending a message to the active user bypass security approvals?

In the current design of Microsoft Copilot Cowork, sending internal communications (Emails or Teams messages) to the user who is currently logged in is not classified as a sensitive action requiring human confirmation. The system assumes that sending data to oneself is inherently safe, failing to account for the fact that these messages can contain malicious triggers or be used to stage data for exfiltration through network requests.

Question: How does an attacker actually get the data out of the Microsoft environment?

Once the agent is manipulated via indirect prompt injection to send a message to the user, the exfiltration occurs when the user opens that message. The message is crafted to trigger attacker-controlled network requests—often through features like URL previews or embedded media—which then transmit the gathered data to an external server controlled by the attacker.

Question: Is this vulnerability limited to a specific AI model?

No. The research indicates that this attack achieved a high success rate against various state-of-the-art models. Specifically, the researchers noted that Claude Opus 4.7 was among the models successfully exploited, indicating that the vulnerability is a result of the system's architectural design and integration with M365 rather than a flaw in a specific LLM.

Related News

US Tech Giants Target Australia for AI Data Center Expansion Amidst 9 Gigawatt Capacity Proposals
Industry News

US Tech Giants Target Australia for AI Data Center Expansion Amidst 9 Gigawatt Capacity Proposals

US technology firms are increasingly identifying Australia as a strategic destination for artificial intelligence data center development. This interest is reflected in a massive pipeline of infrastructure projects, with current proposals reaching a total capacity of 9 gigawatts. However, recent industry data reveals a significant gap between these ambitious plans and their actual realization. As of June, none of the 9 gigawatts of proposed capacity had been commissioned. This suggests that while the intent to expand AI infrastructure in the region is high, the industry is currently navigating a complex transition phase where proposed projects have yet to reach operational status. The situation highlights both the immense potential of the Australian market and the current bottlenecks preventing the immediate deployment of large-scale AI computing power.

The Frontier AEO Tracker: Analyzing Astra Project Trends and Frontier Model Selections for DX Leaders
Industry News

The Frontier AEO Tracker: Analyzing Astra Project Trends and Frontier Model Selections for DX Leaders

Latent Space has officially launched the Frontier AEO Tracker, marking the debut of its inaugural Astra project. This initiative is specifically designed to monitor and analyze Answer Engine Optimization (AEO) trends across leading frontier models, including Astra. Developed in response to high demand from founders and Developer Experience (DX) leaders, the tracker provides critical insights into the selection processes and behaviors of advanced AI systems. By focusing on what frontier models prioritize, the project aims to offer a comprehensive overview of the evolving AI landscape. This tool serves as a strategic resource for stakeholders looking to understand the mechanics of model-driven information retrieval and how to navigate the shifting paradigms of digital discovery in the age of frontier AI.

Decoding the AI Avalanche: A Comprehensive Guide to Opaque Recurrence and Essential Industry Terminology
Industry News

Decoding the AI Avalanche: A Comprehensive Guide to Opaque Recurrence and Essential Industry Terminology

The rapid ascent of artificial intelligence has introduced a significant volume of new terminology, described by industry experts as an "avalanche" of terms and slang. To address this growing complexity, TechCrunch AI has released a specialized glossary curated by Natasha Lomas, Romain Dillet, Kyle Wiggers, and Lucas Ropek. This guide focuses on defining the most critical words and phrases that individuals are likely to encounter in the current technological landscape, including complex concepts such as "opaque recurrence." As the AI field continues to expand, understanding this evolving vocabulary is essential for navigating the technical and social implications of the technology. The glossary serves as a foundational resource for both professionals and enthusiasts attempting to keep pace with the industry's linguistic shifts.