The Short Answer
You can add a paywall and keep your AI citations if the thing you charge for is not the thing assistants quote. Keep the answer readable to everyone, people and crawlers alike, and put the paywall on the action: the download, the export, the full dataset, the tool's saved results. If you must gate the text itself, put the direct answer in the free part at the top of the page and mark the gated section with Google's paywalled content markup so it is not mistaken for cloaking. What loses citations is gating the answer, blocking crawlers to protect it, or hiding it in a way that shows machines one page and people another.
The one sentence version
Gate the download, never the answer, and when you do gate text, keep the answer above the paywall and label the gated part honestly.
Why Paywalls Cost Citations
An assistant quotes what it can read. Training crawlers copy pages, search crawlers index them, and live agents fetch them while a person waits for an answer. If any of those receives a paywall instead of the content, the page stops being a source. It may still rank on its title, and a person may still pay, but the sentences an assistant would have repeated are no longer available to it.
That loss is quiet. Your analytics will not show it, because none of these readers run your scripts, and your search impressions may not move for weeks. The first sign is often that an answer which used to name you now names someone else. If you want to see it in your own data, the live fetches in your server log are the place to look, as our guide on tracking ChatGPT citations explains.
| Who reads the page | What it does with it | What a paywall does to it |
|---|---|---|
| Training crawlers such as GPTBot | Copy the page for future model training | The model learns the paywall message, not your content |
| Search crawlers such as Bingbot, Googlebot and OAI-SearchBot | Index the page so it can be found and fetched later | The page can rank on its title but has little to match a question |
| Live agents such as ChatGPT-User | Fetch the page while a person waits for an answer | The agent gets the paywall and quotes another source |
| People | Read, decide and maybe pay | The only reader a paywall is actually meant for |
Three Kinds of Paywall, Three Different Risks
Not every paywall is the same, and the choice between them decides most of the outcome before you write any code.
| Model | What is free | What is gated | Risk to citations |
|---|---|---|---|
| Action gate | All the content, explanations and examples | The download, export, saved work or full tool output | Lowest. The quotable part stays open |
| Lead in gate | The opening of the page | The rest of the text | Medium. Safe if the answer sits in the free opening |
| Hard gate | A title and a teaser | Everything else | Highest. Assistants have nothing to quote |
The action gate is the one we recommend for most products, and it is the one we used ourselves. It works because of how people search. In six months of search data across 41 sites, queries asking for something to be computed or produced were clicked about five times as often as queries asking for an explanation. The explanation gets summarized in the answer. The action still needs your site. Charge for the action and the answer can stay free to quote.
What We Did on Our Own Product
On the 4th of October 2026 we added a paywall to a document builder we run. Every page stayed public: the builder itself, every template, every guide and every example. Nothing a person or a crawler reads was put behind it. What became paid was the file export, after two free downloads, with a short paid trial for people who want to keep going.
The content stayed open
No guide, template or explanation moved behind the wall, so nothing an assistant could quote changed.
The wall sits on the action
Only the export is counted, and only after two free uses, so a first time visitor always gets something real.
There is an off switch
One setting turns the wall off instantly, so if anything goes wrong it can be reversed in seconds, not in a redeploy.
It was tested before launch
47 checks against the server, 27 checks in the real interface and 11 against production, before the switch was turned on.
It is too early to publish results from a change made days ago, and we will not pretend otherwise. What we can say is why the design looks the way it does: the pages that earn citations and the action that earns revenue were separated before the wall went up, so one could change without touching the other.
If You Gate the Text, Put the Answer Above the Wall
Sometimes the text is the product, for a publisher or a research service. Then the lead in model is the honest compromise, and the rule is simple: the answer goes in the free part. Research on where assistants take their quotes found roughly 44% of citations come from the first 30% of a page, which is exactly where a lead in paywall leaves content open. Our guide to writing for AI extraction covers how to structure that opening.
Write the answer in the first two sentences
State the conclusion, the key number and the date before any background, so the free part stands on its own.
Keep one complete section free
A free section that answers one question fully is worth more to an assistant than a teaser cut mid sentence.
Gate the depth, not the verdict
Method, full data, worked examples and templates are good paid material. The headline finding is not.
Use a real class on the gated container
Wrap the paid section in an element with its own class name, because the markup will point at it.
The Markup That Keeps It Legitimate
Google's documentation is direct about why this matters: the structured data for subscription and paywalled content exists to differentiate paywalled content from cloaking, which violates its spam policies. If crawlers can read text that people cannot, you need to say so in the markup.
The required property is isAccessibleForFree, set to false for a paywalled article. Inside hasPart you describe the gated section as a WebPageElement, with its own isAccessibleForFree set to false and a cssSelector that points at the class on the gated container. A minimal version looks like this:
Minimal paywalled content markup
{ "@type": "Article", "isAccessibleForFree": false, "hasPart": { "@type": "WebPageElement", "isAccessibleForFree": false, "cssSelector": ".paywalled" } }
| Rule from the documentation | What it means in practice |
|---|---|
| Only use class selectors for cssSelector | Point at a class such as paywalled, never at an id or a tag |
| Do not nest content sections | Each gated section is a sibling, not a section inside another |
| Several gated sections go in an array | List each one as its own WebPageElement inside hasPart |
| Make sure Googlebot can access the page | The crawler must be able to reach the full content you are marking up |
| Keep bot access the same on AMP and other pages | If you serve AMP, it must follow the same rules as the main page |
Our wider guide to schema for AI search covers how this sits alongside your other structured data.
Server Side or Client Side
How the paywall is built decides what each reader receives, and it is worth knowing before an assistant finds out for you.
A client side overlay
The full text is in the page and a script covers it. Crawlers that do not run scripts, which is most of them, read everything. That is why the markup above is not optional for this design.
A server side gate
The server only sends the gated text to people who have paid. Crawlers and live agents get the free part. Safer for revenue, and the free part must carry the answer.
Showing bots more than people
Sending the full text only to recognised crawlers while people get a wall is the pattern the markup exists to legitimise. Without it, it looks like cloaking.
Live agents add a twist. OpenAI's documentation says ChatGPT-User visits a page because a user asked, and that robots rules may not apply to it. If your server sends that agent the full text, you have effectively handed the paid content to whoever asked ChatGPT about it. For gated text, treat live agents the way you treat an anonymous visitor. Our guide to JavaScript and AI crawlers explains which agents run scripts and which do not.
Do Not Block Crawlers to Protect the Content
The tempting shortcut is to add every AI crawler to robots.txt and call the content protected. It does protect the text, and it also removes the page from the answers it used to appear in. If the goal is to be paid for access by AI companies rather than to be cited, there are mechanisms built for that, such as charging per crawl with a 402 status code and the RSL licensing standard. Neither is widely honoured by the largest AI companies yet, so decide which outcome you want before you block anything.
Decide what you are selling first
If you sell access to text, gate the text and accept fewer citations. If you sell an action or a tool, keep the text open and gate the action. Most products are the second kind and gate the first by mistake.
How to Test What Each Reader Receives
Fetch the page as an anonymous visitor
Save the response without running scripts and read what is actually in it. That is close to what most crawlers see.
Fetch it again with crawler user agents
Request the page with the user agent strings of Googlebot, Bingbot and ChatGPT-User and compare the three responses with the anonymous one.
Check the answer is in every version
The first two sentences of the answer should appear in every response, paid or not.
Validate the markup
Run the page through the Rich Results Test from Google and confirm the selector matches a real element.
Watch the log for two weeks
Compare live agent fetches and search impressions with the two weeks before the change.
If the anonymous response and the crawler responses differ and you have no paywall markup, fix that first. It is the single most common way a paywall turns into a cloaking problem. Our technical SEO checklist has the wider set of checks to run alongside it.
What to Watch After Launch
| Signal | Where to see it | What a problem looks like |
|---|---|---|
| Live agent fetches | Server log, ChatGPT-User and other live agents | A drop on the gated pages within days |
| Search impressions | Search Console and Bing Webmaster Tools | Gated pages losing impressions for question queries |
| AI impressions | Search Console's AI report and Bing's AI Performance report | Fewer appearances in AI answers for the same pages |
| Paid conversions | Your own checkout data | Revenue flat while citations fall means the wrong thing is gated |
The Bottom Line
A paywall and AI citations can live together, as long as you charge for what people come to do rather than what assistants come to quote. Gate the download, the export or the tool's output and leave the answer open. If the text itself is the product, keep the answer above the wall, mark the gated section so it is not mistaken for cloaking, and treat live agents like anonymous visitors. Then measure both sides, revenue and citations, because a paywall that earns money by quietly removing you from every answer is a trade you should make on purpose, not by accident.
