Paper · Essay · 15 min · February 2026
Entity building: how to craft your identity in Google's Knowledge Graph
In one week Google deleted more than three billion Knowledge Graph entities. What remains is a precision model. The work is to become a node machines can verify, not a string they skip.
In June 2025, Google removed more than three billion entities from its Knowledge Graph in a single week. Kalicube, which has tracked the graph since 2015, put the cut at 6.26 percent of the database: twice what had been added the year before.
It was not a bug. It was a change of model, from accumulation to precision. Fewer nodes, better defined. The reason is practical. AI Overviews, AI Mode and Gemini need a fact base they can trust. A dumping ground of half-typed strings does not serve that.
A brand that sits among the entities Google can name, type and corroborate is protected. A brand that is only a string, with no anchor in the graph, is invisible to a recommendation. Generative systems do not transfer trust to a node they cannot identify.
By 2026 the Knowledge Graph still holds on the order of 1.6 trillion facts about tens of billions of entities. Retrieval has moved. Systems look for knowledge nodes: connected, disambiguated, confirmed by more than one source. The rest of this piece is how to become one of those nodes.
Google's great clarity cleanup: three shifts redefining the Knowledge Graph and its AI future.
01The diagnostic: thirty seconds
Before any markup, measure what the systems already say. Open ChatGPT, Gemini and Perplexity. Ask who the brand is, or who the person is, by name.
Recognised
A detailed, correct reply. Skip to the connections.
Homonym
The name exists. It is attached to the wrong life.
Absent
No information. The node is not there.
The third answer is the one that costs traffic. A recommendation is a transfer of trust. If the system cannot say who produced a page, it will cite the producer it already knows. Content quality does not enter the comparison.
Ahrefs Brand Radar
Enter the brand and three to five competitors. AI Share of Voice asks a simple question: out of a hundred transactional prompts, how often does the brand appear in the top three recommendations? Zero percent means the systems are reading text. They are not reading an entity.
02Four pillars: a file of evidence
Building an entity is not “adding Schema”. It is assembling a convergent body of evidence a machine can check without asking the brand for permission. Kalicube's index covers more than 71 million brands. Search Engine Land reports that Google draws on tens of thousands of sources to corroborate a node. The website is one of them.
1 · JSON-LD
Organization or Person. A declaration of existence, on the page.
2 · Wikidata
The passport into the public graph.
3 · Third parties
Crunchbase, the press, directories, a verified company page. Reputation someone else will sign.
4 · Knowledge Panel
Visible proof that Google has typed the node.
JSON-LD
On the homepage (a brand) or the author page (a person), the markup has to be complete, not cosmetic. The properties most sites omit:
sameAs: LinkedIn, the site, Wikidata, the press profile. Without it, the systems see four people.knowsAbout: the subjects. Do not leave the model to guess.disambiguatingDescription: who this is, and who it is not.@id: a stable identifier so the rest of the site can point at the same node.
Wikidata
Wikidata feeds Google's graph, a large share of chatbot answers, and the Knowledge Panel. Eligibility is less strict than Wikipedia's. A name, a type, a founding date, a website and the social identifiers are enough to exist. Put the Wikidata URL in sameAs. That closes the loop.
Third-party sources
A node that only talks about itself is weak. Crunchbase, a verified LinkedIn company page, a trade directory, a reported interview, Google Scholar, a professional association. Search Engine Land cites a 0.664 correlation between brand mentions and visibility in AI Overviews. Each external mention is a vote.
Knowledge Panel
The panel is not requested. It appears when the graph is sure. Kalicube's Knowledge Graph API Explorer will show a Machine ID (KGID) if one exists. Claim the panel. A correction Google accepts is a correction Gemini will eventually repeat.
03Disambiguation
A generic name, a homonym in another country, a namesake that closed after a scandal: this is the slow failure. During the June 2025 cleanup, Kalicube recorded a 15.27 percent drop in entities typed only as “thing”. Poorly defined, no precise class. Single-type entities rose from 23.9 percent to 28.7 percent. Google is rewarding nodes that are one thing.
- 01
disambiguatingDescription: who this is, who it is not. - 02Wikidata P1889 (“different from”): a formal split that language models ingest as structure, not as prose.
- 03
alternateName: the spellings people actually type.
04The entity gap
An isolated node is a weak node. Power is connections. Search Engine Land's advice is still the right one: build a small internal graph in which each page reinforces a topic the brand actually owns.
In Ahrefs Site Explorer, filter the queries that already trigger AI Overviews. List the adjacent concepts the systems cite and the site does not cover. Each missing concept is an entity gap. For each gap: a dedicated page, marked up, linked from the pages that already rank, with about pointing at the Wikidata item for that concept.
This is not a content calendar. It is topology. The work is to lay the roads that connect the brand to the subjects a machine already treats as neighbours.
05The person as a node
People are entities as strong as brands, especially in health, finance and law, where personal expertise is the filter. Ahrefs Brand Radar shows systems citing named authors. Without those signals, even a well-marked organisation will not be the one recommended.
- An author page as the hub: ProfilePage, sameAs, knowsAbout, hasCredential, worksFor.
- Publish under a real name on other platforms. A guest article, a podcast, a trade appearance is a corroboration link.
- The chain has to hold: Article to author page to sameAs to a verified profile to a person who exists. A break anywhere and the trust score falls.
06What to measure
Traffic and rankings do not describe this work. Some forecasts put a sharp drop in traditional organic volume through 2026 and 2028. Treat those as forecasts. Measure the node directly.
Who is this?
Monthly, on ChatGPT, Gemini, Perplexity. Record how the answers change.
Share of Model
Ahrefs Brand Radar. Out of a hundred transactional prompts, how often in the top three?
Brand search
Brand queries rising while SEO is flat often means a system is already naming you.
Knowledge Panel
Kalicube. A Machine ID means the panel can exist.
When those move and the articles have not been rewritten, the architecture is doing the work. Acquisition cost can fall because the education happens before the first visit. That is the compound effect of a node the systems already trust.
What to remember
The June 2025 cleanup was a message. Accumulation is over. Clarity is the constraint.
Schema makes a page readable. Entity building makes a producer memorable. Once a system knows who you are (identity, expertise, verified links) it stops scoring every new URL in isolation. It starts from a prior. That prior is the advantage. It is also the part of the job that cannot be generated on demand.