Search across our learning center -- articles, newsletters, and more. Start typing or click a topic above.
Stay in the Loop
Get exclusive publishing strategies, industry insights, and early access to new features. No spam -- just signal.
Join 2,000+ publishers. Unsubscribe anytime.
GEO
AI Content Licensing for Publishers: Should You Block AI, License It, or Get Cited in 2026?
Publisher In a Box17 min read
Table of Contents
Your content is already training AI, whether or not anyone paid you for it. You have seen the headlines about the giant licensing checks, and somewhere between the News Corp deal and the next story about a publisher blocking Google, a real question forms: should you block the AI crawlers, try to license to them, or let them cite you and move on? The honest starting point is that the huge checks go to a handful of huge publishers, and the decision that actually matters for everyone else is a different one. This is a monetization and stability question first and a licensing question second, so we will answer it in that order.
AI content licensing is the arrangement where an AI company pays a publisher to use its content, either to train a model or to quote it inside an answer. That is the whole idea. What makes it confusing is that the market has split into two very different things wearing the same name, and most of the advice online talks about the version that will never reach you.
What AI content licensing actually is, and the two ways it pays
There are two payment shapes underneath every deal. The first is the lump-sum training license, where an AI company pays once, or on a multi-year schedule, for the right to train on an archive. The second is pay-per-use, where the publisher gets paid each time its content is crawled or each time it appears inside an answer. The first shape is what makes the news. The second shape is what a smaller publisher can actually plug into, and we will get there.
The training deals are real and they are large. News Corp signed a multi-year license with OpenAI in May 2024 reported to be worth more than $250 million over five years including credits, covering the Wall Street Journal, Barron's, the New York Post, the Times of London and more, according to Variety. Amazon and the New York Times announced a deal in May 2025 that reports later put at roughly $20 million to $25 million a year, the paper's first AI license, per TechCrunch and GeekWire. Reddit licensed its content to Google in early 2024 in a deal reported at about $60 million a year, though by July 2026 Reddit was publicly weighing whether to renew it, according to CNBC. Treat every one of those reported amounts as an estimate rather than a confirmed figure, because almost none of these contracts disclose real terms.
The market underneath them is growing. Grand View Research estimates the AI training-dataset market at roughly $3.2 billion in 2025, rising toward $3.9 billion in 2026, with individual training-data deals ranging anywhere from $5 million to $250 million, a spread Quartz documented across the sector. So the money is genuine. The problem is where it lands.
Reported annual value of named AI licensing deals
USD per year (reported estimates)
Source: Variety, TechCrunch, GeekWire, CNBC. Figures are reported estimates in millions, not company-confirmed, and are ranges rather than guarantees. Every named deal here is a national-scale publisher. That is the pattern, not the exception.
Why the big licensing checks skip most publishers
Look at that chart again and notice what all three have in common. They are national brands with decades of archive and a legal team. AI companies negotiate training deals with a small number of very large rights-holders because that is cheaper and cleaner than paying millions of independent sites. The result is a concentration that several 2026 analyses have named plainly. A Nieman Lab report in May 2026 described the emerging licensing market as a double bind for smaller publishers, where the same large technology platforms that took the content re-emerge as the toll collectors who decide the price. A separate Nieman Lab piece in late 2025 concluded outright that the long tail of publishers will see no meaningful AI licensing revenue, and a Brookings analysis reached a similar verdict about the same gatekeepers building new tollbooths.
That is the part the excited headlines bury, and it matters because it changes your decision. If you are a Digital Publisher running a page, a niche site, or a small portfolio, waiting for a training check is not a plan. It is a lottery ticket. The useful questions are the ones you can actually act on: do you let AI take your content for free, do you charge it or block it at the door, and either way, how do you turn AI attention into revenue you control?
The block-or-license decision, and the trap inside it
Here is the bind the biggest publishers are stuck in, described in their own words, because it tells you exactly what to watch for at your own scale. People Inc, the company formerly known as Dotdash Meredith, told Press Gazette in early August 2026 that it is not blocking Google, but that it keeps the option ready. Its chief executive Neil Vogel said the company is "clearly not turning this off now," and then added the line that matters, "but it is a tool that we can use."
We are clearly not turning this off now. But it is a tool that we can use.
The reason People Inc cannot simply switch AI off is the trap, and it is worth understanding because it applies to you too. Google uses one crawler for both classic search and its AI answers. Block the AI and you block search along with it. Vogel put the timing bluntly, saying the company is "nearly out the other side of search being a material driver of value for us, but we're not there yet." Google now accounts for around 21 percent of the traffic to People Inc's core brands, down from roughly two-thirds at its peak. When a company that large is still measuring whether it can afford to walk away from Google, a smaller publisher has to be even more careful, because search referrals are usually a larger share of a small site's revenue, not a smaller one.
21%
Google's share of traffic to People Inc's core brands in 2026, down from roughly two-thirds at peak
Source: Press Gazette, August 2026
The licensing side has its own trap, and UK publishers named it directly. Press Gazette reported in August 2026 that Google's licensing offers to UK publishers arrive as take-it-or-leave-it two-year packages that bundle the right to train on content with the right to summarize it, and that publishers who call publicly for collective bargaining still sign the individual deals privately for the short-term cash and the referral traffic. One source called it a "no-win prisoner's dilemma." Jason Kint, who runs the trade group Digital Content Next, explained why the platforms prefer it that way, noting that if a company "had to pay everybody for licensing their content, then that affects their margins." So the choice is rarely a clean yes or no. It is a negotiation you usually enter from the weaker side, which is exactly why the infrastructure that lets you set your own terms is the more interesting story.
The infrastructure that lets a smaller publisher set a price
The useful shift in 2026 is that setting terms with AI crawlers no longer requires a boardroom. Several systems now let a publisher of any size allow, block, or charge for AI access, and this is the technical layer worth learning because it maps directly to what PIB calls Technical Retrievability, the machine-readable signals that decide how, and on what terms, an engine can use your pages.
Start with the crawler itself, because the imbalance is the whole argument. Cloudflare measured AI crawl-to-referral ratios in June 2025 and found Anthropic's crawler pulling content at roughly 70,900 times the rate it sent a visitor back, with the company noting that these models "continue to consume more content, more frequently, despite sending the same or less traffic to the source." Read that as the plain economics of the situation. AI takes far more than it returns, so the default of letting it crawl for free is the worst of the three options.
70,900 to 1
How many times Anthropic's crawler pulled content versus sending a referral back, June 2025
Source: Cloudflare Radar. Cloudflare notes app-based traffic can lack referrer data, which can overstate the ratio.
The tools that change that default fall into a few buckets:
Charge at the door
Cloudflare launched pay per crawl on July 1, 2025 as an opt-in control that lets a new domain choose to allow, charge, or block AI crawlers, and a year later, on July 1, 2026, Cloudflare said that from September 15, 2026 it will block the training and agent crawlers by default on the ad-carrying pages of every newly onboarded domain, while leaving the search crawlers allowed. A publisher can block a bot, let it crawl free, or charge it, using the HTTP 402 Payment Required response so a crawler either sends payment intent or gets a price back. Because Cloudflare fronts a large share of the web, this is the single most direct lever a small site has. TollBit runs a similar tollbooth model, charging AI companies each time they fetch a page and letting the publisher set the rate, then keeping most of the revenue for the publisher.
License as a group
The Really Simple Licensing standard, or RSL, launched in September 2025 as an open way to embed licensing and royalty terms directly into the same file that already tells crawlers what they may touch, per Search Engine Land. It supports subscription, pay-per-crawl, and pay-per-inference terms, meaning you can ask to be paid when an AI actually uses your words in an answer, not only when it crawls. Its nonprofit collective, co-founded by an early architect of RSS and a former Ask.com chief executive, is built to license and collect on behalf of the long tail the way a music-rights organization does, so an individual publisher does not have to negotiate alone. Early supporters include Reddit, Quora, Yahoo, Medium, and O'Reilly.
Get paid when the answer uses you
Two models pay on attribution rather than raw crawling. ProRata.ai runs an answer engine built only on licensed content and shares revenue with publishers in proportion to how often their material appears in answers, reportedly with hundreds of publishers signed. Perplexity launched a revenue-share program in October 2025 that pays publishers 80 percent of the subscription revenue from its paid tier, and pays on three signals: human visits, search citations, and what it calls agent actions, where its assistant retrieves an article while doing a task for a user, according to Press Gazette. That third signal is the early shape of a real trend. For the biggest publishers, AI licensing is already a notable and recurring line in earnings: People Inc reported licensing revenue up 26 percent year over year to $40.7 million in Q1, driven mainly by its Meta deal, and Digiday reported from Q1 2026 earnings that publishers are cautiously counting AI licensing as a notable revenue line.
The mechanics under all of this live in two files at the root of your site, and here is the deep-dive most guides skip. Training crawlers and answer crawlers are now separable. You can disallow the training bots in robots.txt while allowing the search-and-cite bots that link back, which lets you refuse to feed a model for free while still earning the citation that sends readers. An llms.txt file guides assistants to your cleanest, best content for quoting, though it is guidance and not a lock, so real enforcement still lives in robots.txt plus a firewall rule for bots that ignore the rules. Wire the paid side in with Cloudflare's controls, a TollBit integration, or an RSL license file, and the decision stops being all-or-nothing. You are no longer choosing between free and blocked. You are setting a price.
The honest play for most Digital Publishers: get cited, then diversify
Now put the pieces together, because the strategy that follows is not the one the licensing headlines imply. For most publishers the training check is out of reach, blocking search to spite AI is a foot-gun, and the tollbooth tools are worth setting up but will not pay the mortgage on their own yet. What is left is the play that has always worked, sharpened for an AI world.
Get cited, then make the citation pay by owning what happens next. When an AI answer names you as the source, that is AI Citation Presence, the share of AI responses in your category that point to you rather than a competitor. It is the closest thing to free distribution left, and it is measurable and improvable. The way you earn it is the same continuous work PIB runs on its own properties: read your own data, find the pages and formats an engine already quotes, and publish more of exactly that, with the human authenticity a model cannot fake. A GEO Readiness Score audit maps where the engines already see you and where they do not, and the fix is the ordinary optimization loop, not a one-time setup.
Then diversify, because that is the real lesson inside the People Inc story. A business that gets a fifth of its traffic from one platform, and cannot turn off that platform's AI without losing its search, is a business hostage to a single channel. The answer is not to win that channel harder. It is to stop depending on it. PIB is a publisher monetization company, a publisher operating system, precisely because a durable publishing business runs across five channels rather than one: Facebook content monetization, Google Discover, content syndication, AI search, and eventually the asset sale itself. Diversification for stability is the whole point. AI licensing, when it does reach you, becomes one more line on that stack, not the thing the business stands on.
That is also why blocking every AI crawler outright is usually the wrong instinct for a growing publisher. The engines are becoming a distribution channel, and a channel you have been cited in is worth more than one you have hidden from. Set your price where the takers give nothing back, keep the door open where they send readers or pay for the answer, and put the energy into the durable work of being the source worth citing across every channel at once. For a fuller map of how the pieces of AI revenue fit together, our guide on how publishers make money from AI walks the three real paths, and our breakdown of the Meta One and AI licensing landscape covers the platform side. If AI answers are already eating your search clicks, what still earns the click in a zero-click world is the companion piece, and GEO, Google Discover and syndication for publishers shows how the channels reinforce each other.
If you want the system rather than the theory, that is what PIB builds. The GEO Authority System runs the LLM Visibility Evaluation and produces your GEO Readiness Score, so you know where the engines cite you today and what moves the number, for $499. The Facebook Automation Machine, the same 75-node workflow PIB runs internally, is the build-it-yourself on-ramp for the distribution and content engine at $397, and the $10K/Mo Profit Playbook maps the revenue path for $197. If you would rather PIB run the pages and the diversification for you, Turnkey Management operates them on revenue-share with no upfront, and Consulting trains your own team to keep 100 percent. The goal is the same either way: a publishing business that survives any single platform changing its mind.
Frequently asked questions
Should I block AI crawlers on my site?
Not indiscriminately. Block the training crawlers that take your content and send nothing back, using robots.txt or a firewall rule, and keep the search-and-cite crawlers that link readers to you. Blocking everything, the way Google bundles its search and AI crawler, can cost you the search traffic you still depend on, which is the exact bind People Inc described in 2026.
Can a small publisher actually get paid for AI licensing?
Directly negotiating a training deal, almost never, because AI companies license from a small number of very large rights-holders. Indirectly, yes and increasingly: Cloudflare pay per crawl, TollBit, the RSL collective license, and revenue-share programs like Perplexity's let a publisher of any size charge for access or get paid on attribution. Expect these to be a supplementary line, not a primary income, in 2026.
What is the difference between a training license and a citation?
A training license pays for the right to feed your content into a model, usually as a lump sum or multi-year deal. A citation is when an AI answer quotes and names you as the source. Training deals are large and rare. Citations are free distribution you can earn and measure through AI Citation Presence, and for most publishers they are the bigger opportunity.
What are robots.txt and llms.txt, and do they stop AI?
Robots.txt tells crawlers which parts of your site they may access, and well-behaved AI bots respect it, so it is where real access control lives. An llms.txt file guides assistants to your best content for quoting, but it is a suggestion rather than an enforcement tool. For bots that ignore both, you need a firewall or a service like Cloudflare that can actually block or charge them.
Is AI licensing worth building a strategy around?
As the whole strategy, no. As one channel inside a diversified publishing business, yes. The publishers who stay stable are the ones who earn AI citations, control AI access on their own terms, and spread revenue across Facebook, Google Discover, syndication, AI search, and asset sales, so that no single platform decision can end the business.
Key takeaways
AI content licensing splits into large training deals, which go almost entirely to national publishers, and pay-per-use access, which any publisher can now set up.
The reported training checks are real, but multiple 2026 analyses conclude the long tail of publishers will see little meaningful licensing revenue.
Blocking AI is rarely clean, because Google uses one crawler for search and AI, so blocking the AI can cost you the search traffic you still need.
New infrastructure, Cloudflare pay per crawl, TollBit, the RSL collective license, and revenue-share programs, lets a smaller publisher charge for access or get paid on attribution.
For most Digital Publishers the bigger opportunity is AI Citation Presence, the free distribution you earn by being the source AI names, measured through a GEO Readiness Score.
The durable answer is diversification for stability, running a publisher operating system across five channels so no single platform can end the business.
Sources
Variety, News Corp Inks OpenAI Licensing Deal Potentially Worth More Than $250 Million (May 22, 2024): https://variety.com/2024/digital/news/news-corp-openai-licensing-deal-1236013734/
TechCrunch, The New York Times and Amazon ink AI licensing deal (May 29, 2025): https://techcrunch.com/2025/05/29/the-new-york-times-and-amazon-ink-ai-licensing-deal/
CNBC, Reddit stock sinks on report it may not renew Google AI content deal (July 22, 2026): https://www.cnbc.com/2026/07/22/reddit-stock-google-ai-content-deal.html
Press Gazette, People Inc not blocking Google 'at the moment' as it rolls out digital subscriptions (August 2026): https://pressgazette.co.uk/news/people-inc-not-blocking-google-at-the-moment-as-it-rolls-out-digital-subscriptions/
Press Gazette, Google gives publishers a 'prisoner's dilemma' (August 7, 2026): https://pressgazettefutureofmediaus.substack.com/p/business-insider-ceo-to-monetise
Nieman Lab, The emerging AI content licensing market puts news publishers in a 'double bind' (May 2026): https://www.niemanlab.org/2026/05/the-emerging-ai-content-licensing-market-puts-news-publishers-in-a-double-bind-a-new-report-warns/
Nieman Lab, Publishers will see no meaningful AI licensing revenue (December 2025): https://www.niemanlab.org/2025/12/publishers-will-see-no-meaningful-ai-licensing-revenue/
Brookings, Same gatekeepers, new tollbooths in the AI content licensing market: https://www.brookings.edu/articles/same-gatekeepers-new-tollbooths-in-the-ai-content-licensing-market/
Cloudflare Blog, The crawl before the fall of referrals (July 1, 2025): https://blog.cloudflare.com/ai-search-crawl-refer-ratio-on-radar/
Cloudflare Blog, Introducing pay per crawl (July 1, 2025): https://blog.cloudflare.com/introducing-pay-per-crawl/
Cloudflare Blog, Your site, your rules: new AI traffic options for all customers (July 1, 2026): https://blog.cloudflare.com/content-independence-day-ai-options/
Search Engine Land, New Really Simple Licensing standard could make AI pay (September 10, 2025): https://searchengineland.com/really-simple-licensing-461834
Press Gazette, Perplexity launches AI subscription revenue-share scheme for publishers (October 2, 2025): https://pressgazette.co.uk/news/perplexity-launches-ai-subscription-revenue-share-scheme-for-publishers/
Digiday, Publishers cautiously count AI licensing as notable revenue in Q1 earnings: https://digiday.com/media/media-briefing-publishers-cautiously-count-ai-licensing-as-notable-revenue-amid-programmatic-strain-in-q1-earnings/
Grand View Research, AI Training Dataset Market: https://www.grandviewresearch.com/industry-analysis/ai-training-dataset-market
Quartz, The price of AI training data, from $5M to $250M: https://qz.com/ai-training-data-pricing-licensing-deals-market-052126
Forbes, These Startups Are Making Sure AI Companies Pay Up For Taking Content (December 23, 2024): https://www.forbes.com/sites/rashishrivastava/2024/12/23/these-startups-are-making-sure-ai-companies-pay-up-for-taking-content/
Written by
Publisher in a Box
The team behind 300M+ managed followers. We help publishers scale traffic, revenue, and audience across Facebook, Google Discover, and syndication networks.