Article 13 min read
Semantic Silos: What Google Stores About Your Site's Focus
by Daniel Sócrates published
You published. You published plenty, and you published carefully. Your competitor published less and shows up more.
The industry always gives the same answer: you lack authority. But authority, put that way, tells you nothing to do on Monday.
There is a better explanation, and it carries the name of a field inside Google.
What the industry calls a silo, and what the machine stores
Silo is the name the industry gave to a simple idea. Group the pages that cover the same subject, link them to one another, and keep distant subjects out of the same corner of the site.
The idea is right. The trouble is that it gets sold as a filing preference, as if it were tidying a drawer. And whoever hears it then asks, fairly: does Google care?
It cares enough to store two numbers about it.
In 2024, a leak exposed the internal documentation of a Google API, the Content Warehouse.
Inside it sits a module whose name gives away the subject: QualityAuthorityTopicEmbeddings.
Authority, topic, and the way the machine represents meaning.
Two fields in that module matter to anyone who publishes. One measures how much your site talks about a single subject. The other measures how far each page drifts from that subject.
This is called site focus, and the distance from it.
How Google measures the focus of a site
The field that says how much a site talks about one subject
The first field is called siteFocusScore. The description in the documentation is short and
blunt: “Number denoting how much a site is focused on one topic”. 1
Look at what that sentence assumes. It assumes a site has a subject. And that you can measure how tightly it concentrates there.
Look at the level, too. The measure is not of the page. It covers the whole site. 1
The field that measures the distance between page and site
Picture your site as a point on a map. Every page you publish is another point. If the pages cover the same subject, they land close together. If one of them covers something else, it lands far away.
The second field measures exactly that distance. It is called siteRadius. The description
reads: “The measure of how far page_embeddings deviate from the site_embedding”. 2
The same module stores the two points that calculation needs: siteEmbedding and
pageEmbedding. It stores a version field as well. Having the three of them together is what
supports the map reading. The site is one point, the page is another, and the distance between
them sits in storage. 3
The name of the module is a checkable fact. That the engineering team put authority, topic, and embedding in the same place suggests intent, and that reading is mine, not the documentation’s. 4
A module mix-up going around the industry
Industry material about site focus tends to cite a third field alongside the two above. It is
called site2vecEmbeddingEncoded, it does exist, and it is being cited in the wrong place.
site2vecEmbeddingEncoded does not belong to the focus module. It lives in another one,
QualityNsrNsrData. 5
Inside that same QualityNsrNsrData sits a field that confuses things further,
siteQualityStddev. It measures dispersion, so it looks like a cousin of the radius. But its
dispersion is of quality, and the one in siteRadius is of topic. Two different
questions, two different fields, two different modules. 6
Checking which module each field came from sounds like fussiness. It is not. Anyone who cites the wrong field is making a claim about topical focus with a measure that says nothing about topic.
What these fields do not tell you
This part matters more than the previous one, and it almost never shows up.
The leak shows what is stored. It does not show how much it weighs. None of these fields comes with a percentage, a public scale, or a disclosed target. One of the fields in the module is a version number, which indicates that the values change from version to version. And all of it is a snapshot of 2024 architecture. 7
So we know Google stores a measure of focus. We do not know how much it decides.
There is more, and this part has an immediate practical consequence. No Google API,
dashboard, or report exposes siteFocusScore or siteRadius. None. 8
So when a tool sells you a “site focus score”, it is handing you its own estimate, calculated by its own method. It may be a good estimate. It simply is not Google’s number, because Google’s number never leaves Google.
Anyone promising otherwise is selling what they do not have.
Two more proofs that the judgment covers the whole site
One leaked document, on its own, is a single proof. There are two others, and both are public.
The policy that speaks of signals a site accumulates
Google keeps a policy on site reputation abuse. Its definition says, in official text: “Site reputation abuse is a tactic where third-party content is published on a host site mainly because of that host’s already-established ranking signals, which it has earned primarily from its first-party content”. 9
Read the second half of that sentence again. Google states that a site accumulates ranking signals, and that it earns them primarily from its own content.
The policy is about hosted third-party content, and that is all it is about. What matters here is what it declares in passing: something accumulates at the site level, and it comes from what the site itself publishes.
The timeline shows the matter was taken seriously. Announced on March 5, 2024, alongside a core update. In force on May 5, 2024. First manual actions reported from May 7. And an update on November 19, 2024, making clear that involvement or oversight by the site owner excuses no one. 10
The 2022 statement, and what became of it
The third proof is older. In August 2022, Google announced the helpful content system. It described the system as a site-wide signal. And it stated that content from sites carrying a lot of unhelpful material tends to perform worse in search. 11
Here comes the part nobody tells you alongside it. In March 2024, Google announced that the system had been folded into its core ranking systems. I went to check the current official page on August 1, 2026: that wording is no longer there. 12
The statement still holds as a statement from 2022. Four years of distance from today, and a warning that the text has moved.
None of the three documents mentions the other two. Reading the three as pointing at the same thing, a judgment that happens at the site level, is mine. The facts in each one are checkable. The stitching between them is interpretation. 13
What was measured from the outside
Breadth of coverage weighed more than domain authority
On March 23, 2026, Kevin Indig published a study in Growth Memo on how artificial intelligence systems pick the sources they cite. The base was more than 21,000 citations. 14
Two findings matter to anyone mapping territory. First, breadth of topical coverage weighed more than domain authority in source selection. Second, in low-concentration sectors, a focused strategy of 30 to 50 pages competes. 14
One boundary, and it comes from the study itself. That work measures source selection by AI. It does not measure ranking in Google, and it measures none of the fields in the earlier section of this article. The two things talk to each other, and they are not the same thing. 14
What page size has to do with it, and what it does not
The same study measured page size against citation. Pages above 20,000 characters averaged 10.18 citations. Below 500 characters, 2.39. The largest single jump appears between 5,000 and 10,000 characters, and the effect has a ceiling. 15
I put that number here for a specific reason. Someone will ask whether the way out is writing giant pages, and the data itself replies. What got measured was a correlation between size and citation by AI, with a ceiling declared in the study. Reading it as a character-count instruction invents a recipe the data does not support. 15
Cover your own territory before fighting for someone else’s
Breadth and depth are two strategies
I measured the query tail of two properties in my own portfolio, over the same window, from May 1 to July 30, 2026. Both are Brazilian sites, measured in Search Console.
In the first, the ten largest queries add up to 10.8% of the observable universe, spread over 101,328 distinct queries. In the second, the ten largest add up to 42.7%, and the hundred largest reach 83.5%, over 3,750 queries. 16
The first site spreads. The second concentrates. I call both of them legitimate map strategies, and that reading is mine: the same number that signals dilution in one case signals coverage in the other. What decides is the business, not the ruler. 16
Where the growth actually showed up
In a third case, in the health sector, I broke the growth down at the address level. The sample had 128 addresses and 2,332 clicks. This one is a Brazilian site as well. 17
Of that total, 56.5% came from pages that did not exist three months earlier. The other 43.5% came from old pages gaining position. Three old pages multiplied their own clicks by 27, by 41, and by 34. 17
Publishing new work grows a site. But almost half the growth came from pages that were already sitting there. They started being read another way once the rest of the territory was covered.
No client is identified in this article. Sector kept, name removed, by an editorial rule that holds for every piece this firm publishes. 18
Your Monday exercise
Ten minutes, no tools at all.
Open a sheet of paper and list the subjects your site covers today. One per line. Next to each one, write how many pages you have about it.
Now look at the list and answer this: which of those subjects is yours? If more than three fight for the title, your site has a blurred center. Every new page gets measured against an average that represents nothing.
When you want to see this drawn instead of imagined, the next step is to crawl your own site and look at the link map. The how is in How to find errors on your site with Screaming Frog. This article is the why.
Frequently asked questions
Is a semantic silo the same thing as a blog category?
Not necessarily. A category is a label you pick in a dashboard. Focus is a measure that comes out of comparing what your pages say. You can have ten tidy categories and a blurred center, and you can have zero categories and high focus. What counts is what the pages say, and how much they say the same thing. 1 2
Can I see my siteFocusScore in some tool?
No. No Google API, dashboard, or report exposes that field or siteRadius. Any number sold
under that name is an estimate from the tool that calculates it, by its own method. 8
How many pages does a subject need?
The Indig study, from March 2026, found that in low-concentration sectors a focused strategy of 30 to 50 pages competes. The boundary holds: that work measures source selection by AI, and not ranking in Google. 14
Do bigger pages get cited more often?
On average in that study, yes, with a ceiling. Pages above 20,000 characters had 10.18 citations against 2.39 for pages under 500. It is a correlation measured in citation by AI, not a size instruction. 15
Notes and sources
siteFocusScore, moduleGoogleApi.ContentWarehouse.V1.Model.QualityAuthorityTopicEmbeddingsVersionedItemin the documentation of the Content Warehouse API leaked in 2024. Literal description quoted. Live verification on the public mirror of the documentation on August 1, 2026.siteRadius, same module, with the literal description quoted. Live verification on August 1, 2026.siteEmbedding,pageEmbeddingandversionId, same module.- Name of the module
QualityAuthorityTopicEmbeddings. The name is a checkable fact; the reading of engineering intent is the author’s inference. site2vecEmbeddingEncodedbelongs toGoogleApi.ContentWarehouse.V1.Model.QualityNsrNsrData, not to the focus module. Live verification on August 1, 2026.siteQualityStddev, inQualityNsrNsrData, measures quality dispersion, distinct from the topical dispersion insiteRadius.- The leak shows what is stored and never the weight. No field carries a percentage, a public scale, or a disclosed target. Snapshot of 2024 architecture.
- No Google API, dashboard, or report exposes
siteFocusScoreorsiteRadius. - Google Search Central, “Spam policies for Google web search”. Literal quotation verified live on August 1, 2026.
- Timeline of the same policy: announced on March 5, 2024, in force on May 5, 2024, first manual actions from May 7, 2024, update on November 19, 2024. Sources: Google Search Central Blog, posts of March 5 and November 19, 2024. The literal text of those two posts was not checked in direct reading and therefore does not appear in quotation marks.
- Google Search Central Blog, announcement of the helpful content system, August 2022. The page did not open for literal checking during reporting, and the content is reported indirectly, without quotation marks.
- In March 2024 Google announced that the helpful content system had become part of its core ranking systems. The current official page, consulted on August 1, 2026, no longer contains that wording. Negative verification.
- The convergence of the three proofs is the author’s inference. None of the three documents mentions the other two.
- Kevin Indig, “The science of how AI picks its sources”, Growth Memo, March 23, 2026. Base of more than 21,000 citations. Boundary declared by the study itself: it measures source selection by AI, not ranking in Google.
- Same study as note 14: pages above 20,000 characters averaging 10.18 citations against 2.39 for pages below 500 characters, largest jump between 5,000 and 10,000 characters, effect with a ceiling.
- Search Console, query dimension, two properties in the author’s portfolio, window of May 1 to July 30, 2026. Both are Brazilian sites. Denominator declared as the observable universe of queries. The reading of breadth against depth as two legitimate strategies is the author’s.
- Growth decomposition at the address level, sample of 128 addresses and 2,332 clicks, same portfolio, Brazilian site.
- Anonymization: sector kept, identification removed. No client name, domain, property identifier, or page address is reproduced.
Let us look at what the machine understands about your company
A diagnostic conversation, with no slide deck. You bring your website address. We bring what Google and the AI assistants already know about it today.