TL;DR: On May 6, 2026, Canada's Privacy Commissioner and three provincial counterparts published the results of a joint investigation into OpenAI's ChatGPT. The verdict: OpenAI violated federal privacy law (PIPEDA) and provincial laws in Quebec, British Columbia, and Alberta. The company scraped vast amounts of personal data from the public internet (including health information, political views, and children's data) without obtaining consent. It launched ChatGPT knowing the model fabricated facts about real people. It ran without data retention or deletion policies. And until April 2024, it used a deceptive design pattern that forced users to give up their chat history to opt out of training data use. OpenAI disagreed with the findings but agreed to corrective measures on a 3-to-6-month timeline, including quarterly compliance reports. Two provinces, BC and Alberta, refused to fully resolve the matter, saying consent for scraped data simply cannot be obtained after the fact.
What the Commissioners Found
The investigation started with a simple question: did OpenAI follow Canadian privacy law when it built ChatGPT? The answer, across all four jurisdictions, was no [1].
Commissioner Philippe Dufresne and his counterparts in Quebec, British Columbia, and Alberta spent months examining how OpenAI assembled its training data for GPT-3.5 and GPT-4. What they found was a company that moved fast and broke privacy law in the process [1][2].
Five specific violations stood out:
- Overcollection: OpenAI gathered "vast amounts of personal information without adequate safeguards," including sensitive data like health conditions, political views, and information about children [3]
- No consent: The company scraped publicly accessible websites and social media platforms without obtaining valid consent from the people whose data it took [2][3]
- Known inaccuracy: An OpenAI co-founder admitted the model "likes to fabricate things." The company shipped it anyway, without establishing any accuracy baseline for outputs about real people [4]
- No deletion policies: OpenAI ran without formal retention or disposal rules for the personal data it collected. Datasets sat around "as long as necessary to train successive iterations," which could mean forever [3][4]
- Restricted access rights: Canadians lacked any effective way to access, correct, or delete personal information OpenAI had collected about them [3]
The Consent Problem That Can't Be Fixed
Here's the detail that should worry every AI company watching this case.
OpenAI's web crawler, GPTBot, didn't launch until August 2023 [4]. That means website owners had no mechanism to opt out during the critical data-gathering phase. By the time the crawler existed, the training data had already been collected. The opt-out tool arrived after the barn door was open, the horse was gone, and someone had built a billion-dollar product from the hay.
British Columbia and Alberta's privacy commissioners took the hardest line on this point. Their position: consent for scraped personal data simply cannot be obtained retroactively. The data was taken without permission, and no amount of post-hoc compliance measures can fix that [1][4].
Both provinces found the complaint "well-founded but unresolved," meaning they accepted OpenAI's corrective measures as steps forward but refused to say the underlying violation was fixed [1].
Quebec went further, issuing province-specific recommendations and declaring the consent issues unresolved under its own statute [1].
The Deceptive Design Pattern
For users who actually signed up for ChatGPT, things weren't much better.
At account creation, OpenAI showed a single notification: "Don't share sensitive info. Chats may be reviewed and used to train our models" [4]. That was it. One line. For a service that would store, analyze, and train on every conversation you had with it.
It got worse. Until April 2024, any user who wanted to opt out of having their conversations used for training had to sacrifice access to their chat history. Want privacy? Lose your data. The investigators called this a "deceptive design pattern," the polite regulatory term for a dark pattern [4].
OpenAI eventually decoupled these settings. But for the first 18 months of ChatGPT's existence (the period of explosive growth from zero to hundreds of millions of users) opting out came with a punishment.
What OpenAI Agreed To
Despite disagreeing with the findings, OpenAI committed to a set of corrective measures on a staggered timeline [1][3]:
- Immediately: Publish a Canadian blog post explaining its privacy practices and confirm deprecation of GPT-3.5 and GPT-4
- Within 3 months: Expand its model development documentation with plain-language explanations of how training data is used, and add a notice to the signed-out ChatGPT experience explaining training data use
- Within 6 months: Improve personal information accessibility in its data export tool, finalize retention controls for deprecated datasets, test protections for minor family members of public figures, and begin quarterly compliance reports to regulators
The federal commissioner deemed the complaint "well-founded and conditionally resolved," the regulatory equivalent of "you broke the law, we believe you'll fix it, and we'll be watching" [1].
What This Actually Means
Canada is the first country to complete a formal privacy investigation into ChatGPT and publish enforceable corrective measures. Italy temporarily banned ChatGPT in 2023 over similar concerns. The EU's data protection authorities have been investigating since early 2023. But Canada got to the finish line first [1][2].
The findings matter beyond Canada's borders because they establish a clear regulatory position: scraping the public internet for AI training data without consent violates privacy law. Full stop. Not "might violate" or "raises concerns about." Violates.
Every AI company that trained models on Common Crawl data, Reddit posts, social media profiles, or public forums just got put on notice. If Canadian regulators can find OpenAI in violation, the same logic applies to every competitor doing the same thing.
The BC and Alberta position is particularly sharp: you cannot fix a consent violation after the fact. If you didn't have permission when you took the data, having permission now doesn't help. That principle, if adopted by other jurisdictions, would require AI companies to either obtain consent before training or accept permanent legal exposure for their existing models.
The Bigger Picture
Commissioner Dufresne used the announcement to push for something larger: modernizing Canada's federal privacy law for the AI era [1]. Canada's Privacy Act and PIPEDA were written before large language models existed. The investigation showed they're still powerful enough to find violations, but enforcement remains limited to compliance agreements rather than fines.
Compare that to the EU, where GDPR violations can cost up to 4% of global annual revenue. For OpenAI, that could mean billions. Canada's regulators found the same violations but lack the same teeth.
Meanwhile, OpenAI faces a separate court order requiring it to hand over 20 million anonymized ChatGPT conversations to The New York Times as part of a copyright lawsuit [4]. The company is being squeezed from multiple directions (privacy regulators, copyright holders, and its own users) all asking the same question: what did you do with our data?
Sources
- Office of the Privacy Commissioner of Canada: Joint investigation into OpenAI's ChatGPT leads to better protections (May 6, 2026)
- Canada's National Observer: OpenAI did not respect Canadian privacy laws in developing ChatGPT, probe finds (May 6, 2026)
- Office of the Privacy Commissioner of Canada, Background: Summary of joint investigation into OpenAI's ChatGPT (May 6, 2026)
- PPC Land: Canadian regulators find ChatGPT privacy rules broken from the start (May 2026)
Published: May 7, 2026