AI & SEO

The Entity Before
the Keyword

Fernando Angulo
Senior Market Research Manager at Semrush, an Adobe company
9 Min Read
July 29, 2026

Three people share this name. The machine has to pick one.

There are three people named Fernando Angulo with their own entry in Wikidata. One of them is a Spanish basketball player. One is a Peruvian ornithologist. One is me. Before any AI engine decides whether to cite my research on AI search, it has to decide which of us it is talking about. That decision happens before relevance, before quality, before any keyword I have ever written. Most brands have the same problem and are still optimizing one layer too late.


Quick Answer:

Entity Density is the degree to which a machine can resolve who or what you are without guessing. It is built by three practices: one canonical description string used identically across every surface, one machine-readable anchor that ties your identifiers to a single resolvable node, and corroboration from the third-party reference sources generative engines over-cite.

“Things, Not Strings” Turns 14

The idea that you should optimize for entities rather than keywords is not a discovery of the AI era. It has a precise public birthday. On 16 May 2012, Amit Singhal announced Google’s Knowledge Graph in a post whose subtitle became the most quoted line in the history of semantic search: “things, not strings.” Singhal described it as “an intelligent model, in geek-speak, a ‘graph’, that understands real-world entities and their relationships to one another.” It shipped that day with more than 500 million objects and more than 3.5 billion facts about and relationships between them.

That was 14 years ago. I want to be exact about this, because a lot of current writing on entity optimization presents it as something generative AI invented in the last eighteen months, and that framing is both wrong and strategically useless. It is wrong because the mechanism predates ChatGPT by a decade. It is useless because it implies the discipline is immature and you have time, when the opposite is true: this is a mature idea with a mature toolset that most organizations simply declined to implement while it stayed optional.

I apply this correction to my own work as a rule. The GEO framework I teach was named by Aggarwal et al. at Princeton and Georgia Tech in November 2023, not by me. Answer Engine Optimization was set out by Jason Barnard in 2018, not in 2024. Entity optimization belongs to Google’s 2012 announcement and to the semantic search practitioners who built on it. Getting the lineage right is not humility, it is accuracy, and accuracy is the entire product when your job is research.

What is genuinely new is narrower, and it is the subject of the rest of this post: the cost of getting your entity wrong has changed by an order of magnitude.

What Changed Is the Cost of Ambiguity

Generative engine optimization (GEO), the practice of getting your brand mentioned and cited inside AI-generated answers rather than ranked among ten blue links, runs on a different pipeline than classic search. The engine does not hand back a list and let a human disambiguate. It resolves the entity itself, then retrieves against that resolution, then writes a single answer. If it resolves you incorrectly, there is no second result for the user to click instead. You are simply absent, or worse, described as someone else.

The scale of that filtering is measurable. Semrush’s AI Visibility Index 2026, built on 126 million US AI-search prompts gathered between January and April 2026 across 22 verticals and four platforms, found that an answer carries a hard ceiling on how many sources it will use: roughly 15.4 cited sources per answer on ChatGPT, 9.2 on Google AI Overviews, and just 3.3 on Gemini. On a Gemini answer, three slots decide the entire visible universe for that question. Ambiguity is not a soft ranking penalty in that environment. It is elimination in the first round.

The Index also separates two things most teams still report as one number. Being mentioned in an answer and being cited as its source diverge sharply by platform: the overlap between the brands an answer names and the domains it actually links runs from about 64% on Google AI Overviews down to 30% on Gemini. I have written about why that gap is the real KPI in citation authority and AI visibility. The entity layer sits underneath both. A machine cannot cite a node it cannot identify.

Two more findings from the same dataset point the same direction. Only 36 global brands held top-100 visibility across all four platforms for the full study period, a concentration that gets more extreme by sector: in News and Media the top three brands take 82.9% of all visibility, against 41.4% in Finance. And the engines lean heavily on reference platforms when resolving who someone is: Semrush’s own summary of the Index notes that ChatGPT “frequently relies on community and reference platforms such as Reddit and Wikipedia.” Those are entity databases. The engines are not consulting them for prose style. (Research data referenced here is © Semrush.)

I Failed This Test on My Own Entity

Here is the part that made me write this post rather than another framework explainer.

My own site has carried a disambiguation instruction for over a year, in the visible copy, in the disambiguatingDescription field of my schema, and in the llms.txt file that exists specifically to tell language models who I am. It said, in effect: this is the Fernando Angulo at Semrush who researches AI search, not the Ecuadorian footballer and not the Colombian boxer.

While researching this article I did the obvious thing and checked. Searching Wikidata for my exact name returns three people. There is Fernando Angulo the Spanish basketball player (Q5859182), born 11 June 1967 in Miranda de Ebro, a small forward who played for Baskonia, Fuenlabrada and Tenerife between 1988 and 2007. There is Fernando Angulo the Peruvian ornithologist (Q48815491), affiliated with CORBIDI. And there is me (Q138975073).

No Ecuadorian footballer. No Colombian boxer. For a year I had been carefully instructing every machine that reads my site that I am not two people who do not exist, while the two who do went unmentioned. And the Peruvian ornithologist is almost certainly the source of a nationality error that kept surfacing in AI answers about me, one I had already spent time correcting in my bios without ever finding its cause.

I have corrected it across the site. I am including the failure because it is more instructive than the fix: I write about this discipline professionally, I had done the work, and the work was aimed at the wrong targets because I had never verified the target list against a primary source. Disambiguation asserted from memory is not disambiguation. It is noise that looks like diligence.

Entity Density: The Three Practices

What you are building is not a page. It is a node the machine can resolve confidently. I call the property Entity Density: the degree to which a machine can resolve who or what you are without guessing. Three practices build it, in this order.

1. One canonical string, used identically everywhere. Pick the exact description of who you are and paste that same string, character for character, into every profile you control: LinkedIn, Crunchbase, X, YouTube, GitHub, conference speaker bios, podcast show notes. Not variants. Not “Semrush (Adobe)” in one place and “SEMrush” in another. Entity resolvers merge signals by matching tokens, so every variant you publish is a vote to split your node into two weaker ones. This is unglamorous, it is mostly copy-paste, and it is the practice teams skip because it does not feel like strategy.

2. One machine-readable anchor. A canonical string tells a machine what you claim. An anchor gives it somewhere to resolve that claim to. In practice this means a Wikidata item where one is warranted, a stable @id in your Person or Organization schema, and a sameAs array that lists every profile from practice one. The sameAs array is the single highest-leverage line in most schema blocks, and it is the one most often left empty. Structuring it correctly is its own topic, which I covered in schema markup for AI search.

3. Corroboration from sources the engines over-cite. Your own claims about yourself are the weakest evidence in the system, because every entity makes them. What moves an engine is independent confirmation from the reference sources it already leans on, which is why the Index singles out Wikipedia and Reddit as platforms ChatGPT returns to repeatedly. This is the slowest of the three practices and the only one you cannot execute unilaterally, which is exactly why it carries the most weight.

The order matters. Corroboration built on an inconsistent canonical string produces corroboration for a node that is not yours. Fix the string, plant the anchor, then earn the references.

What Entity Density Doesn’t Fix

Three honest limits, because a framework that claims to solve everything is a sales pitch.

It does not replace having something worth citing. Entity Density determines whether a machine can identify you; it says nothing about whether your material deserves a citation slot. Perfect resolution of a node with nothing behind it just makes your irrelevance unambiguous.

It is slow, and it is slowest exactly where it matters most. Practices one and two you can complete in a week. Practice three runs on other people’s publishing schedules and can take quarters. Anyone selling entity work with a 30-day guarantee is selling practices one and two and calling it the whole job.

And it does not transfer across languages by itself. An entity resolved cleanly in English can be unresolved or wrongly merged in Spanish, because the corroborating source pool is different and thinner. That asymmetry is a real opening rather than just a problem, which is the argument in Cross-Lingual Citation Authority.

One closing note on where the demand actually sits. In Semrush’s US keyword data pulled for this article in July 2026, “semantic seo” runs about 6,600 searches a month, “knowledge graph seo” 1,300, “entity seo” 1,000, and “entity based seo” 480 at a $9.31 CPC, the highest commercial intent in the set. But the sharpest signal is in the question data: the variations of “how do I find entities” add up to roughly 730 searches a month, more than the definitional queries they compete with. People have accepted the premise. They are asking for the method. That gap, between a settled idea and an unexecuted one, is where the 14 years went.

Frequently Asked Questions

An entity is a thing rather than a string: a specific person, company, product, or concept that a search system holds as a distinct node with its own identifier and relationships. Google introduced the idea publicly on 16 May 2012 when it launched the Knowledge Graph under the line “things, not strings,” shipping with more than 500 million objects and 3.5 billion facts about them. Optimizing for an entity means making that node resolvable and well-described, rather than repeating a phrase on a page.

Start with your own before you map related ones. Search your exact brand or personal name on Wikidata and see how many distinct items already carry it, that set is who the machine has to choose between. Then ask each major AI platform who you are and read what comes back. If the answers disagree with each other, you have an entity problem, not a content problem, and no amount of keyword work fixes it.

They overlap but are not the same. Semantic SEO is the broader practice of covering a topic and its related concepts in a way a machine can model. Entity optimization is narrower and more identity-focused: making sure a specific node, your brand or your person, is unambiguous, correctly described, and corroborated. You can have strong topical coverage and still be an entity the engine cannot resolve.

Traditional SEO optimizes a page against a query. Entity optimization optimizes a node against the rest of the web, and most of the work happens off your own site: consistent descriptions on third-party profiles, a machine-readable identifier, and corroboration from reference sources. That is why a site can rank well and still be invisible or misdescribed inside AI answers.

The public starting point is 16 May 2012, when Amit Singhal announced Google’s Knowledge Graph with the phrase “things, not strings.” Entity-based optimization has therefore been available for about 14 years. What changed recently is not the idea but the penalty for ignoring it: generative engines resolve an entity before they decide whether to cite it, so ambiguity now costs visibility rather than just a knowledge panel.

Bring the entity layer to your stage

I keynote on AI search, GEO, and citation authority across the US and Europe, including how entity resolution decides who gets cited before relevance is even considered.

Invite me to speak → Download the GEO framework
Recommended Reading

Latest Insights

View all articles