AI Training Data Theft: Permission Won't Fix It | Elacity
Unsealed filings show Microsoft named AI training the largest theft of labor. Copyright notices were stripped and paywalls bypassed. Permission never protected creators; ownership at the source can.
Microsoft Called AI Training the Largest Theft of Labor. Permission Was Never Going to Stop It.
The essays, photos, and code you published are already inside models you will never be paid by. Complaints about AI training data theft are old news. What changed this month is that the company doing it wrote down, in its own internal memos, that theft was the accurate word.
The Admission Behind AI Training Data Theft
In filings unsealed this month in The New York Times case against OpenAI and Microsoft, a Microsoft executive called training on unlicensed work the largest theft of labor in human history in an internal memo. Another passage in the record described it as an astonishing theft of unprecedented proportions. According to the same unsealed material, the companies bypassed paywalls and stripped copyright notices from the work they trained on. One OpenAI leader wrote that the products are largely substitutive for the very work they learned from.
Read that last admission slowly. The output competes with the input, the input was taken without payment, and the people taking it understood both facts. This was not a gray area someone wandered into. It was a plan, described in the words of the people running it.
Permission Was the Thing That Failed
Notice what protected none of those creators. Copyright notices did not, because the filings say they were stripped. Paywalls did not, because they were bypassed. Terms of service did not, because nobody scraping at that scale reads them. Every one of those is a permission: a note attached to a file that asks the reader to behave. A note can be removed. A file lying in the clear cannot defend itself.
The courts are not rushing in either. The Justice Department has argued that training on copyrighted text is fair use. Anthropic paid 1.5 billion dollars to settle a related claim, but that was for keeping pirated copies, not for training, and settlements land years after a model has already learned. One running tracker now counts well over 140 US copyright suits against AI companies, and not a single one un-trains a model. Litigation can price the past. It cannot protect the next file you publish.
Some licensing deals are real, and some creators are finally being paid. That is genuine progress and worth stating plainly. A license, though, is still permission with an invoice attached. It works only where the counterparty chooses to ask, and the unsealed memos are a record of what a large company does when it decides not to.
A File That Cannot Be Taken in the Clear
There is another place to put the boundary, and it is not a courtroom. Put it inside the file. Elacity dDRM wraps a song, a dataset, a model, or a document into a Wealth Capsule: an encrypted, programmable good that carries its rights and royalties inside it. The work stays sealed everywhere except the split second it is used, and even in that instant the key that unlocks it is used, never handed over. The secret exists in the clear only for a moment, inside a sealed sandbox, welded to one transaction, then wiped.
The key itself is split across an owned quorum of independent machines, and each one re-checks your on-chain rights before it releases its share. No single operator, Elacity included, holds the whole key. A copyright notice can be stripped because it sits next to the file. A rule that lives inside the encryption cannot be peeled off, because there is no clear file to peel it from. If a model wants the work, it meets the terms at the point of use, or it gets nothing to train on.
This is what the Microsoft memo was really admitting. The industry did not defeat your rights. It walked around them, because the work was lying in the open. Technology encodes a theory of power. A file anyone can copy and relabel encodes one answer about who owns your labor. A file that pays you when it is used encodes another. That second answer is the core of what we mean by Sovereign Capital: productive property you hold, not a claim you file after the fact.
The Honest Edge
Be precise about what this does and does not do. It does not reach into a model that already trained on your open work, because that data is gone, and pretending otherwise would be its own dishonesty. It is trust-minimized, not trustless: a quorum that colluded could in principle rebuild a key, which is exactly why the quorum is owned and auditable rather than one company's promise. Today the forensic watermark that traces a leak back to a buyer covers images, not yet video or audio, and the consumer portal for wrapping and selling your own work is still being built. The deeper argument, that value now flows to owning what the machine needs rather than to selling it your time, we made in Universal Basic Equity. The practical version, wrapping your work so it pays you, is how to turn your data into capital.
The memo gave this era its honest name. The response cannot be another note asking to be respected. It has to be property that enforces its own terms, so the next thing you make is something a machine pays to use rather than something it quietly absorbs. You were the product. Now you can own the asset class. Follow Elacity on X for how the ownership layer gets built.