When Vendors Retrain on Your Data Without Asking Permission

The Quiet Rewrite: How Vendors Are Claiming Your Data for AI Training

Atlassian - Data Contribution - your organisations content - no blanket permission
Atlassian - Data Contribution - your organisations content - no blanket permission
Photo: Angus Fox — via admin.atlassian.comCC BY 4.0

Across enterprise software, a structural shift is underway. Vendors that once positioned themselves as custodians of customer data are quietly rewriting the terms of that relationship, amending service agreements to permit the use of customer-generated content for AI model training. The mechanism is rarely dramatic. It arrives as a terms-of-service update, or buried in a changelog or communicated through a notification most users dismiss without reading.

TL;DR – You should demand that software agreements include an explicit AI training opt-out that is active by default, not buried in settings or dependent on a user locating a consent toggle.

Atlassian offers a documented example of this pattern. The company introduced opt-out defaults for certain AI-related data practices, meaning customers who took no action are automatically enrolled. For enterprise teams managing sensitive project data, legal documentation, or proprietary product roadmaps inside tools like Jira or Confluence, the burden of protection shifts from the vendor to the customer. Opting out requires awareness, deliberate action, and in many cases, administrative access that not every stakeholder holds.

You can see that I turned 'Data Contribution' to 'Off' and left Atlassian the following feedback. As an administrator it would be unlikely that I could give permission for my organisations data to be used by anyone 'to improve apps for everyeon' I think it is bonkers in terms of over reach as a permission. 

I cannot responsibly grant blanket permission for all end-user data to be used for AI training. I have no visibility into which models or downstream services may receive the data, who may ultimately have access to it, which jurisdictions it may be processed or stored in, how long it will be retained, or how future uses may change. I also cannot be certain whether contributed data contains PII, commercially confidential information, privileged communications, or data subject to contractual or regulatory obligations. A single global on/off switch is not an adequate governance or risk management control for enterprise data. I need granular controls over what is shared, for what purpose, with whom, for how long, and with auditable transparency.

The pattern extends well beyond productivity software. Investigations across multiple jurisdictions have scrutinized Meta's AI-enabled smart glasses for capturing and processing data from bystanders who provided no consent whatsoever, meaningful or otherwise. Reports also claim that outsourced workers were able to view sensitive content including nudity and private interactions filmed by the company's AI smart glasses. These cases illustrate a broader dynamic: the collection perimeter is expanding, while the consent framework is contracting.

What emerges is a recognizable form of function creep. SaaS platforms adopted for document management, customer support, or project collaboration are being repositioned, without explicit customer agreement, as AI training infrastructure. The data flows remain invisible to most users, and the commercial benefit accrues entirely to the vendor. For compliance officers and procurement teams, this represents a material shift in counterparty risk that existing vendor contracts were never designed to address.

Regulatory Exposure: GDPR Purpose Limitation and the Secondary Use Problem

When a vendor retrains its AI models on data that customers submitted for an entirely different operational purpose, it triggers one of the most consequential obligations under European data protection law. Article 5(1)(b) of the GDPR enshrines the principle of purpose limitation, requiring that personal data be collected for specified, explicit, and legitimate purposes and not further processed in a manner incompatible with those original purposes. Using customer interaction data, query logs, or uploaded documents to improve a vendor's foundational model is, in almost every realistic scenario, precisely that kind of incompatible secondary use.

The PII and Privileged Data Compounding Factor

The regulatory exposure deepens considerably when the data contributed to vendor systems contains personally identifiable information, legally privileged communications, or commercially sensitive material. Employees routinely paste client records, contract drafts, financial projections, and internal strategy documents into AI-assisted workflows without any awareness that this content may later be ingested as training data. Once such material enters a training pipeline, the organization that originally submitted it loses meaningful control, and the vendor may lack the technical means to isolate and delete specific contributions upon request, creating direct tension with GDPR's data subject rights under Articles 17 and 21.

Dark Patterns and the Collapse of Informed Consent

Compounding these risks are the consent mechanisms vendors deploy, which frequently rely on dark patterns to obscure opt-out controls. Consent to model training may be bundled within lengthy terms-of-service updates, pre-ticked by default, or buried behind multiple settings screens. The effect is that enterprise customers nominally agree to secondary use without any genuine understanding of what that agreement entails. Under GDPR, consent must be freely given, specific, informed, and unambiguous. Consent obtained through deliberately opaque interfaces does not meet that standard, exposing both the vendor and the data controller to regulatory scrutiny.

Organisations should treat undisclosed vendor retraining not merely as a contractual breach but as a live compliance risk requiring immediate legal assessment, data mapping, and, where necessary, supervisory authority notification.

What Organisations Are Actually Risking: Confidentiality, Sovereignty, and Competitive Harm

Atlassian - Data Contribution - your organisations content - Survey if you say no
Atlassian - Data Contribution - your organisations content - Survey if you say no
Photo: Angus Fox — via admin.atlassian.comCC BY 4.0

When vendor contracts permit AI training on client-submitted data, the consequences extend well beyond abstract privacy concerns. Organisations face a layered set of exposures that touch on confidentiality, legal jurisdiction, and long-term competitive position. Understanding each dimension is essential for any enterprise evaluating its AI vendor relationships.

Confidentiality Risks in Training Pipelines

Sensitive enterprise data - including legal documents, financial records, internal communications, and strategic plans - routinely passes through AI-powered tools. Once that data enters a vendor's training pipeline, the organisation loses meaningful control over how it is processed, retained, or reproduced. As TrustArc's research highlights, AI systems can inadvertently regurgitate information on which they were trained, creating a direct pathway for confidential material to surface in outputs delivered to unrelated third parties. This risk is compounded by what researchers term perpetual processing: AI systems often retain information in ways that make it difficult to trace or erase, undermining standard data lifecycle controls.

Data Sovereignty and Jurisdictional Exposure

Training pipelines frequently operate across multiple jurisdictions, meaning data submitted in one regulatory environment may be processed under an entirely different legal framework. For organisations subject to sector-specific obligations - in financial services, healthcare, or public administration - this creates direct compliance exposure. Cross-border training activity can place data outside the reach of the originating jurisdiction's enforcement mechanisms, eroding the sovereignty guarantees that many regulatory regimes are designed to protect.

Competitive Harm Through Purpose Drift

Perhaps the least visible risk is competitive. Proprietary workflows, pricing logic, client segmentation strategies, and operational metadata can all function as training signals, even when the organisation believes it is simply using a productivity tool. This reflects a well-documented pattern of purpose drift, where data collected for one reason is repurposed in ways the originating organisation never anticipated or authorised. Once embedded in a shared model, these signals may indirectly benefit competitors using the same platform.

Taken together, these risks demand that organisations move beyond reviewing headline privacy commitments and instead scrutinise the specific contractual provisions governing secondary data use before deployment.

Governance and Contractual Safeguards Organisations Must Demand Now

Atlassian - Data Contribution - your organisations content - On by default
Atlassian - Data Contribution - your organisations content - On by default
Photo: Angus Fox — via admin.atlassian.comCC BY 4.0

Reactive responses to vendor data misuse are insufficient. Organisations must embed enforceable protections into procurement and governance frameworks before contracts are signed, not after confidential data has already been ingested into a third-party model.

Contractual Clauses That Must Be Non-Negotiable

Every SaaS and AI vendor agreement should contain explicit, unambiguous language across three critical dimensions. First, data minimisation obligations must restrict vendors to collecting only what is strictly necessary for the contracted service. Second, defined retention limits should specify maximum storage periods and mandate verifiable deletion upon contract termination. Third, and most critically, agreements must include an explicit AI training opt-out that is active by default, not buried in settings or dependent on a user locating a consent toggle. Purpose limitation clauses should further prohibit secondary use of organisational data for any model improvement, benchmarking, or product development without separate, documented consent.

Vendor Assessment Criteria

During procurement, legal and information security teams should evaluate vendors against the following criteria:

  • Least privilege architecture: systems should access only the data required for each discrete function
  • Granular consent controls: administrators must be able to configure data-use permissions at a feature level, not solely at account level
  • Transparent audit logs: vendors should provide documented records of how, when, and where customer data is processed
  • Clear policy change notification: contractual obligation to provide advance written notice before any amendment to AI training terms

Internal Governance Steps to Implement Immediately

Contractual protections are only effective when paired with robust internal governance. Organisations should prioritise three actions. First, audit all existing SaaS agreements to identify contracts that lack explicit AI training restrictions, treating these as material risk exposures requiring renegotiation. Second, conduct AI-specific privacy impact assessments for every tool that processes personal, sensitive, or commercially confidential data, a practice endorsed by leading privacy frameworks and increasingly expected by regulators. Third, establish ongoing monitoring protocols that track vendor policy updates, because terms of service can change with limited notice, and silent amendments represent one of the most common vectors for function creep.

Organisations that embed these safeguards into both supplier relationships and internal governance cycles are best positioned to retain meaningful data sovereignty in an environment where vendor incentives and organisational interests are not always aligned.