Something that has been redefined in value over the past few years is located on the archive floor of a London media firm. It is now made up of rows of hard drives instead of filing cabinets, but it still bears the weight of decades of verified reporting. It was always referred to as the archive. It used to refer to cost centers, storage issues, and difficulties with digital preservation. In the pitch meetings, it’s now referred to as training data. It turns out that training data has actual financial value for the corporations creating the massive language models that the media is both secretly negotiating with and scared of.
AI developers are increasingly obtaining direct licenses from British publishers for their archives. Thru an industry group called SPUR, the BBC, the Guardian, the Financial Times, and a coalition of other major UK outlets have been working to establish commercial licensing frameworks. This is done in part to generate revenue and in part to establish the principle that their content has a price before the UK government moves to waive that price thru broad text-and-data mining exceptions, which publishers have been vehemently opposing. The publishers are attempting to win both the policy question and the business question at the same time.

Understanding the financial reasoning is simple and doesn’t require a strong interest in AI. For the past fifteen years, British journalism has witnessed the shift in advertising revenue from print to digital, and then from digital publishers to Google and Meta, who control the majority of online ad expenditure while producing comparatively little of the attention-generating media. Newsrooms have experimented with all possible business model variations, reduced staff, consolidated, and moved behind paywalls. The archive was an asset that wasn’t making money, despite the fact that most publishers have invested a significant amount of money digitizing and preserving it. That is sometimes drastically altered by licensing it to an AI business.
AI developers are finding it most difficult to reject the quality argument. Large language models that are mostly trained on unfiltered web content experience hallucinations, uneven factual correctness, and unreliable sourcing. In contrast to most web content, the material in a major newspaper’s archive has undergone editing, fact-checking, and structuring. It has editorial accountability traces, bylines, and datelines. With uniform standards, it covers decades of recorded history. Access to high-quality journalistic archives is more than just a compliance need for an AI corporation attempting to decrease the frequency with which its model confidently declares something incorrect. This provides publishers with a non-defensive negotiation stance.
The aspect of this story that transcends individual business agreements is the SPUR coalition. Instead of negotiating outlet by outlet in a way that favors the much larger, much better-resourced AI companies in each individual transaction, major UK publishers are working together to establish norms around pricing, attribution, use restrictions, and audit rights that shape how AI companies access British journalism as a category. The analogy to music rights is flawed but instructive: before creating the kind of collective pressure that led to more equal streaming economics, the music industry allowed digital platforms to set low-cost standards for music access for years. Publishers are making an effort not to repeat that order.
When these trades are covered, the attribution issue is not given enough attention. When an AI system uses their content to produce a response, publishers are negotiating not just for remuneration but also for the presentation of that content. For companies whose worth largely depends on confidence in their editorial identity, a model that delivers Financial Times analysis without any indication of its source or integrates Guardian articles without giving credit to the Guardian raises concerns about brand integrity. Although it is still really unknown how strong those provisions are in practice and whether they can be enforced at the volume and speed that AI answers operate, the deals being constructed are making an effort to address this.
