A short, furious essay from msd (quailblog) holding up two cases of mass downloading side by side: Aaron Swartz, co-creator of RSS, and Meta’s AI training.

  • Swartz downloaded about 70 gigabytes of academic articles from JSTOR for the purpose of archiving and sharing knowledge. He was charged so aggressively — up to 35 years in prison, a $1 million fine, and asset forfeiture — that he took his own life rather than face the court fight and financial ruin.
  • Meta, by contrast, torrented over 80 terabytes of books to train its AI models. Its disclosed consequence is a lawsuit it will most likely settle for a fraction of what the models earn.

The author’s framing: Swartz’s use case was the dissemination and archival of knowledge; Meta’s is powering proprietary models that enrich billionaires. Same act of copying, radically different treatment — one treated as a crime worth destroying a person over, the other as a cost of doing business.

A commenter adds the obvious follow-up: Anthropic paid $1.5 billion for scanning books to train its models — real money, but trivial relative to what those models return.

It’s an opinion piece with a chip on its shoulder rather than a measured analysis, but it captures a double standard that keeps showing up in AI’s training-data story.