Insolvency estates as AI training data
When a company is wound down, its complete document estate (years of correspondence, contracts, invoices and ledgers) passes to an insolvency administrator. Under a written mandate from that administrator, the estate can be rights-reviewed, anonymized and licensed as AI training data. This is TGDC's founding sourcing channel.
- Source
- Document estates of wound-down European companies
- Authorization
- Written administrator mandate before a single file moves
- Typical span
- Decades of correspondence, contracts, ledgers and forms
What a document estate contains
A single wound-down company leaves behind the full paper trail of its operating life: email correspondence with customers and suppliers, signed contracts and their amendments, invoices and dunning letters, ledgers and closing statements, HR files, and the procedural documents of the winding-down itself. Unlike a web scrape, an estate is coherent. The same counterparties, deals and disputes thread through thousands of documents over years. That is exactly the longitudinal structure that document AI needs to learn, and no public dataset carries it.
The legal path
- Administrator mandate. The insolvency administrator controls the estate's assets, including its data. Licensing proceeds only under a written mandate from the administrator.
- Rights review. Before anything ships, each corpus goes through a review covering ownership, third-party material and statutory limits: what may be licensed at all, and under which conditions.
- Anonymization. Named entities, personal data and commercial identifiers are removed or deterministically pseudonymized, so cross-document references stay coherent without exposing real parties. Every batch is verified, and a verification report ships with the delivery.
The result is a corpus that is both legal and useful. Ownership is settled by the administrator's mandate, and the rights review sets out what may be licensed from the estate. The anonymization report covers privacy, and it is written for every batch. All of this is on paper before anything reaches a buyer.
Why this data matters for AI
Everything publicly crawlable is already inside every frontier model. Estate data never appeared online, so it is genuinely unseen. That is valuable in two distinct roles. As training material, it adds signal instead of repetition. As evaluation and RL-environment substrate, it is free of the contamination that would otherwise invalidate the results. And because it is real operational work rather than curated text, it carries the noise, ambiguity and edge cases that make trained behavior robust in deployment.
What it looks like as a product
TGDC delivers estates as licensed corpora: archives with documented chain of custody, structured extracts against a customer schema, and held-out evaluation sets. Each delivery includes written chain of title from the administrator mandate onward and the per-batch anonymization report. Buyers can trace every artifact back to an authorization. That is the traceability EU AI Act documentation duties increasingly demand.
Administrators and advisors: if you manage estates and want to understand what a data mandate looks like, write to hello@thegeneraldata.com.