Information Gain SEO: What Google Documents, and What the Field Infers
Share
Information gain SEO is the practice of making a page carry information the pages already ranking do not — and the fact that reorders the whole topic is that Google's own patent scores it against the reader, not against the web. “Contextual estimation of link information gain” (US12013887B2, granted 18 June 2024, assignee Google LLC) defines an information gain score as one “indicative of additional information that is included in the document beyond information contained in documents that were previously viewed by the user”. That last clause is doing enormous work. It makes novelty a property of a reading session rather than a fixed grade on your URL. Meanwhile Google's guide to Search ranking systems — which names BERT, RankBrain, passage ranking, PageRank, freshness, deduplication and a dozen more — never mentions information gain at all. Both of those are true at once, and holding both is the whole discipline.
What information gain means in SEO
Used by practitioners, the phrase means one thing: how much a page adds to a topic beyond what the results already ranking for that query say. A page with high information gain tells a reader something the first page of results does not. A page with low information gain restates the consensus in fresh sentences.
The reason the idea got urgent is not that Google announced anything. It is that answer engines changed the economics of restating. When a generative answer can compress the agreement of ten sources into one paragraph, a page whose only contribution is that agreement has been made redundant by the summary that quotes it. That is the pressure behind the current wave of information gain SEO advice, and the pressure is real even where the mechanism is unproven — see generative engine optimization for how citation behaviour differs from ranking behaviour.
So the concept is worth taking seriously. The specific claims made about it mostly are not, and separating those two is the job of this page.
Three different things share the name
Most confusion here is lexical rather than technical. Three distinct objects wear the same label, and arguments about information gain usually turn out to be two people talking about two of them.
- The statistics term. In machine learning, information gain is the reduction in entropy achieved by splitting on a feature — the standard criterion for choosing splits in a decision tree. It is decades old, has nothing to do with search, and is what you will find if you look up what is information gain in a textbook rather than a marketing blog.
- The Google patent. A specific, readable, granted document that borrows the term for document ranking and defines it precisely. It exists, it is public, and almost nobody quoting it has read past its title.
- The SEO practice. The working advice — add something new — which is sound guidance whose stated justification is usually the patent, and whose actual justification is that it is simply how you write something worth reading.
Keeping the three apart costs nothing and immediately dissolves questions like “what information gain score should I target”, which is a category error: the statistics term has no target, the patent publishes no threshold, and the practice is not scored. Most disagreements about information gain seo dissolve the moment both parties say which of the three they mean.
What the Google information gain patent actually says
The google information gain patent people cite is “Contextual estimation of link information gain”. The most recent grant in the family is US12013887B2 — inventors Victor Carbune and Pedro Gonnet Anders, assignee Google LLC, priority date 18 October 2018, granted 18 June 2024, and still listed as active. Two earlier grants, US11354342B2 and US11720613B2, share the title.
Its own abstract is short enough to quote rather than paraphrase. It describes “determining an information gain score for one or more documents of interest to the user”, where that score “is indicative of additional information that is included in the document beyond information contained in documents that were previously viewed by the user”, and says the score “may be determined for one or more documents by applying data from the documents across a machine learning model”. Documents are then presented “in a manner that reflects the likely information gain that can be attained by the user”.
Read plainly, that is a re-ranking mechanism for a reader who is working through a topic, not a quality grade stamped on a URL. It is also, and this matters, a patent. Google files thousands. A granted patent proves an idea was novel enough to protect and worth the fee; it does not prove the idea shipped, survived, or ever touched web search rankings.
Documented, versus widely believed
The divergence here is unusually wide, which makes it worth setting out as a table rather than prose. Every row below is checkable against a primary source in one click.
| Claim | Status | Primary source |
|---|---|---|
| Google holds a patent describing an information gain score | Documented | US12013887B2, granted 18 Jun 2024 |
| The score measures what a document adds beyond documents the user already viewed | Documented | US12013887B2 abstract |
| Google reduces near-identical pages in results | Documented | Ranking systems guide — deduplication systems |
| Google works to show original reporting ahead of those who cite it | Documented | Ranking systems guide — original content systems |
| Information gain is a named Google Search ranking system | Not documented | Absent from the ranking systems guide |
| There is an information gain score, threshold or percentage to hit | Not documented | No published figure exists |
| Adding N original facts per page raises rankings | Inference | No source; reasonable practice, unproven mechanism |
Sources: US12013887B2, Google Patents · A guide to Google Search ranking systems, Google Search Central (last updated 10 December 2025).
Notice what the middle rows do. Two mechanisms that behave a great deal like the thing people describe — deduplication and original content systems — are documented, by name, on Google's own page. The behaviour practitioners observe is real and has a published explanation. It just is not the explanation being cited. If you want the documented version of the argument, those two rows are it — and they are the closest thing to published seo content quality signals that Google offers on this subject.
The per-reader twist almost everyone drops
Return to the patent's wording: information gain relative to “documents that were previously viewed by the user”. If a system worked that way, three consequences follow, and none of them appear in typical information gain SEO advice.
- Your page has no single score. It has one per reader state. High for someone arriving cold, low for someone who has just read three articles making the same points — the same page, the same words, a different number.
- Position in a reading sequence matters as much as content. Being the fourth page someone opens is a different job from being the first, and the fourth page's advantage is covering what the first three omitted.
- You cannot optimise a per-session variable directly. You can only make your page's contribution unusual enough that it survives most plausible reading histories — which, conveniently, is exactly the same instruction as “have something to say”.
That is why the practical advice is right even though the mechanism is unproven. The action a per-reader score would reward, the actions deduplication and original content systems demonstrably reward, and the action that earns a citation from an answer engine are all the same action.
Information gain is not uniqueness
The single most expensive mistake in this area is treating information gain content as a plagiarism problem. It is not. Unique content seo checks answer a textual question: do these exact strings appear elsewhere? Information gain asks a semantic one: does this meaning appear elsewhere?
A rewritten digest of the top ten results passes every duplication check ever built and adds nothing whatsoever. It is one hundred per cent unique and zero per cent new, and it is the exact failure mode information gain seo exists to name. Original content seo advice that stops at the uniqueness check is therefore measuring the wrong axis — and it is the axis most content tooling reports on, because textual overlap is cheap to compute and informational novelty is not.
The related trap is semantic coverage: matching the entity and subtopic profile of the pages that rank. Useful for relevance, and directly counterproductive here, because perfect coverage of what the field already says is a precise description of a page with nothing to add.
How to actually add information gain
Four things reliably survive the subtractive test. All four are expensive, which is why the shortcut versions of information gain seo — spin the top ten, add a statistic, bolt on an opinion paragraph — do not work.
- Primary sources, read and quoted. Almost everything written about a documented topic is a summary of a summary. Reading the source and quoting its own words — as this page does with the patent abstract — produces material the field's paraphrases do not contain. It is also the cheapest of the four.
- Your own measurements. Numbers from work you actually did cannot be restated by anyone who did not do it. This is the strongest and slowest form.
- A documented-versus-believed correction. Where the field asserts something a primary source contradicts, saying so with the citation attached is high-value and rare, because it requires checking rather than aggregating.
- A method precise enough to repeat. Most guides describe an outcome. A procedure with its thresholds, failure modes and abort conditions written down is a different artefact.
Then apply the subtractive test before publishing: delete every sentence the first page of results already contains, and read what is left. If what remains is a paragraph, you have a paragraph's worth of contribution and should publish a paragraph. If nothing remains, the honest verdict is hold, not build — the same verdict a content gap analysis reaches when demand exists but you have nothing to add to it.
What you can and cannot measure
There is no information gain metric to report, which is the most awkward fact about information gain seo as a practice. Nothing in Search Console exposes one, no third-party tool can compute the patent's version without your reader's browsing history, and any tool offering an “information gain score” is scoring its own invented proxy — usually semantic distance from the top results, which is a reasonable proxy and not the thing itself.
What you can measure is downstream and honest: whether pages carrying original material earn citations in AI answers where restatements do not, whether they attract links that summaries never attract, and whether they hold position through core updates. Those are lagging, noisy, and real — unlike a score. Pair the question with E-E-A-T, which is also not a score and is also frequently sold as one, and with content freshness, where the same temptation to fake a signal rather than earn it shows up in a different costume.
We publish our own position on this rather than hedge it: we do not claim information gain is a live ranking system, because Google does not document one. We build for it anyway, because every documented mechanism that is named rewards the same behaviour. That is the reasoning our Moat X-Ray module applies when it scores how much of a page an AI could reconstruct without it.
Questions we hear
What is information gain?
In machine learning, information gain is a standard measure of how much a feature reduces uncertainty about an outcome — it long predates search. In SEO the phrase means something narrower: how much a page adds beyond what the pages already ranking say. Google's own patent defines it a third way again, as a score “indicative of additional information that is included in the document beyond information contained in documents that were previously viewed by the user”. Three different objects share one label, which is most of why the topic is confusing.
Is information gain a Google ranking factor?
Not a documented one. Google's guide to Search ranking systems names BERT, RankBrain, neural matching, passage ranking, PageRank and link analysis, freshness, original content, deduplication, site diversity, reviews and spam detection — and does not name information gain anywhere. A patent is evidence that an idea was worth protecting, not evidence that it runs in production. Treat every specific claim about thresholds, percentages or scores as inference.
What is the Google information gain patent?
“Contextual estimation of link information gain”, US12013887B2, assigned to Google LLC, invented by Victor Carbune and Pedro Gonnet Anders, priority date 18 October 2018 and granted 18 June 2024. It is one grant in a family that also includes US11354342B2 and US11720613B2. It describes applying a machine learning model to produce an information gain score, then ordering what a user is shown by it.
Does the patent score pages or readers?
Readers, in effect. The patent measures a document against “documents that were previously viewed by the user”, which makes the score a function of one person's reading history rather than a fixed property of your page. The same page scores high for someone who has read nothing on the topic and low for someone who has just read three articles saying the same thing. Most information gain SEO advice quietly drops this and treats the score as static.
How do you increase information gain on a page?
Publish something the field cannot restate: your own measurements, a primary source read and quoted rather than summarised at second hand, a documented-versus-believed correction, or a method described precisely enough to be repeated. The test is subtractive — remove every sentence the first page of results already contains, and whatever survives is the page's actual contribution. If nothing survives, more words will not help, and no amount of information gain seo tooling will change that.
Is information gain the same as unique content?
No, and conflating them is the common mistake. Unique content SEO checks are about textual duplication — whether these exact words appear elsewhere. Information gain is about informational novelty — whether this meaning appears elsewhere. A page can be one hundred per cent unique by a plagiarism checker and carry no new information at all, which is precisely what a rewritten summary of the top ten results is.
Go deeper
Related: SEO content writing · content gap analysis · content optimization · semantic keywords · generative engine optimization.
Want the defensibility read done for you — which of your pages an AI could rewrite from the field's consensus, and which carry something it could not? See how Moat X-Ray works as a module, or request a free AI-powered SEO audit as a first read.
Last updated 23 August 2026. Sources: “Contextual estimation of link information gain”, US12013887B2 (Google Patents, granted 18 June 2024); “A guide to Google Search ranking systems”, Google Search Central, last updated 10 December 2025.
Keep exploring this topic
This guide is part of our Content cluster.