Start free trial of Lex HR →

OpenAI crawler accessed uniVersa customer data, insurer says

UniVersa says an OpenAI web crawler scraped customer records from a misconfigured server on 7 July; the insurer has informed Bavarian data-protection authorities.

28 July 2026

OpenAI's web crawler accessed customer records stored on a misconfigured server run by German insurer uniVersa, the company has told regulators. UniVersa reported the incident to the Bavarian data protection authority after discovering the exposure, raising immediate questions about what information was copied and whether any of it entered model-training datasets.

Heise and Versicherungsjournal reported on 23–24 July that the access occurred during an IT change on 7 July, when a public-facing server was left reachable with insufficient access controls. The reporting says the crawler downloaded records from that server; uniVersa notified the Bavarian data protection authority following discovery. The insurer has not published a full breakdown of which customer fields were exposed or how many records were affected.

OpenAI has not publicly confirmed that any specific uniVersa data were retained or used to train its models. The question of dataset provenance is now central: regulators and privacy experts are focused on whether data captured by automated scrapers are copied into the large, persistent corpora companies use for model development, or whether they are transiently accessed and discarded. That distinction matters under the EU's General Data Protection Regulation, which requires controllers to document lawful bases for processing and implement safeguards for sensitive personal data.

The episode sits against a tightening regulatory backdrop. The European Data Protection Board's recent guidance on web scraping and AI training datasets highlights that automated collection from the open web can still trigger GDPR obligations when scraped material contains personal data. The guidance sets out factors regulators will consider when assessing whether scraping and subsequent model training are compliant, such as the controller/processor roles, the purpose of processing, and the feasibility of technical measures to prevent inclusion of personal data.

For employers and HR teams, the incident is a reminder that third-party scraping incidents can surface through unexpected channels. Insurers and other companies that host customer portals, benefits systems or personnel records on externally reachable infrastructure must ensure configuration management and change-control processes prevent accidental public exposure. Where vendors or external services are involved, contracts should address obligations around preventing and responding to automated harvesting of personal data.

UniVersa's notification to the Bavarian authority is the clearest step mapped to law so far, but several material details remain undisclosed. The insurer has not specified the categories of customer data taken, whether any records contained special categories of personal data, whether affected individuals have been notified, or what remediation steps were implemented beyond reporting the incident. OpenAI likewise has not said whether it archived any of the accessed files or whether they appear in training datasets; there is no public confirmation of an independent audit or third-party forensic review.

The lack of transparency on scope and retention will shape the regulatory and reputational follow-up. The Bavarian authority could require uniVersa to provide a detailed incident report and impose measures or fines if it concludes the insurer failed to implement appropriate technical and organisational safeguards. Separately, questions about dataset provenance and model training could prompt further scrutiny of AI providers' ingestion and retention practices across the EU.

As employers expand use of AI, the uniVersa episode underlines two linked pressures: operational hygiene on public-facing systems, and contractual and technical guarantees from AI suppliers about what their crawlers collect and keep. For HR leaders charged with protecting employee and customer data, the incident reinforces the need to treat web-facing infrastructure as a frontline compliance issue and to insist on clear, auditable commitments from AI vendors about data use and deletion.

Sources
  1. uniVersa: OpenAI AI crawler accessed customer data
  2. Datenschutzvorfall bei uniVersa: Crawler von OpenAI griff Daten ab