Six signals appear to determine citation: extractability, comparative framing, third-party corroboration, entity clarity, recency and specificity match.
An important caveat before the list. These systems are opaque and change without notice. What follows is inferred from observing what gets cited in practice, not from published documentation. Treat them as useful working heuristics rather than confirmed ranking factors and be sceptical of anyone, including this post, presenting them with more certainty than that.
What makes the framework useful is not precision about mechanics. It is that each signal has a different cost, timeline and difficulty, which lets you sequence the work sensibly.
Signal 1: Extractability
What it is: Whether a specific, attributable claim can be lifted from your page without surrounding context.
Why it matters?: Model assembling an answer needs a self-contained statement. If your first extractable claim sits 800 words down behind three paragraphs of context, it may not be reached or a clearer page gets used instead.
What fails?: “We pride ourselves on rapid deployment.” No number, no boundary, no condition.
What works: “Implementation takes four to six weeks for teams under 50 users.”
Cost: An afternoon for five pages.
Timeline: 8–14 weeks to show effect.
Difficulty: low.
This is the cheapest signal to improve and usually produces the largest single gain relative to effort.
Signal 2: Comparative framing
What it is: Whether content positions you relative to alternatives, or only describes you in isolation.
Why it matters: Buyer queries are comparative best tool for X, compare A and B, alternatives to C. A page describing only your product contains nothing that can answer them.
What it requires: Comparison pages against named competitors, alternatives pages, and honest statements about where competitors serve someone better.
The blocker is usually organisational, not technical. Companies avoid naming competitors. The comparison happens regardless buyers ask for it, third parties publish it, models synthesise it from whatever exists. Absence means no input into the framing, not absence of the comparison.
Cost:3–4 hours per page, four to six pages.
Timeline: 8–14 weeks.
Difficulty: low technically, high politically.
Signal 3: Third-party corroboration
What it is: How many independent sources describe you, and how consistently.
Why it matters: A vendor’s claim about itself is one self-interested data point. Fifteen practitioners independently recommending the same tool is materially stronger evidence. These systems look for consensus.
Where it comes from: Review platforms, third-party category roundups, community discussion, independent comparison articles.
Why this is the hardest signal: You cannot manufacture it quickly, and attempts to do so are detectable. It requires sustained outreach, review programmes and genuine community participation over months.
- Cost: Ongoing, 2–3 hours monthly.
- Timeline: 4–6 months minimum.
- Difficulty: High.
This is the signal most companies skip, which is precisely why it remains a source of advantage.
Signal 4: Entity clarity:
What it is: Whether the system holds a coherent representation of what your brand is — category, function, target customer, price point, competitors.
Why it matters: If a model believes you are a CRM when you are a sales engagement platform, you are considered for CRM queries where you compare badly and excluded from your actual category entirely.
What breaks it: Describing yourself one way while G2, Crunchbase and third-party articles describe you another. External consensus generally wins.
The fix: One canonical description, deployed verbatim everywhere, plus Organization and SoftwareApplication schema with a complete `sameAs` array
- Cost: One day.
- Timeline: 4–8 weeks.
- Difficulty: low.
Do this before building visibility. An incorrect entity means you build presence in the wrong category, which is worse than building none.
Signal 5: Recency:
What it is: How current the information about you is, across your own content and third-party sources.
Why it matters: Software changes. A source from 2021 describes a product that may not exist in that form. Systems assessing current suitability weight recent sources more heavily — and engines relying on live retrieval, like Perplexity, weight it more than those relying on trained knowledge.
Where it shows: Review velocity, dates on your content, currency of third-party articles describing you, freshness of comparison pages.
The specific failure: A review profile that stopped two years ago suggests a product that may have stopped too.
- Cost: 15 minutes monthly for review requests, quarterly page reviews.
- Timeline: 4–6 months for review velocity to accumulate.
- Difficulty: low effort, requires consistency.
Signal 6: Specificity match;
What it is: how closely your content matches the narrowness of the query.
Why it matters: The broader the query, the more it favours accumulated presence. The narrower it gets, the more it favours whoever addressed exactly that situation.
Where the advantage sits: “best project management software” favours whoever is mentioned everywhere. “Project management for architecture firms under 50 people” favours whoever has a page addressing precisely that — and frequently nobody does.
Why this matters more in AI search than in traditional search: prompts are longer and more situational than search queries. People type fragments into Google and describe their circumstances to an AI.
- Cost: 2–3 hours per use-case page, four to eight pages.
- Timeline: 8–14 weeks.
- Difficulty: low.
This is the structural advantage available to smaller companies.
What order should you work on these?
Sequence by cost and dependency, not by importance.
- Order Signal Why here?
- Entity clarity Everything else depends on being categorised correctly
- Extractability Cheapest, fastest, largest gain per hour
- Comparative framing Highest-value content, finite build
- Specificity match Where a smaller company can actually win
- Recency Low effort, needs to start early because it accumulates
- Third-party corroboration Slowest start in month two, lands in month six
Two things about this ordering:
Entity clarity comes first not because it is most important but because it is a dependency. Building visibility while miscategorised means building it in the wrong category.
Third-party corroboration comes last in sequence but should start early, because its lead time is the longest. Begin outreach in month two even though nothing has moved outreach in week seven produces coverage in month five.
What is not on this list?
Four things frequently cited as signals that do not appear to function as such:
Keyword density: These systems are not matching keyword frequency.
Word count: Claim density matters; length does not.
Backlink volume: Links help pages rank and be discovered, which affects availability for retrieval. But the direct signal is corroboration across independent sources, not link count.
llms.txt: No confirmed evidence major systems read it. It costs an hour and does no harm, and it is not a lever.
And one honest omission: product quality. These systems cannot assess it. They observe how products are described, which is a proxy that fails entirely when nobody has described you.
Frequently asked questions:
Are these confirmed ranking factors?
No. They are inferred from observing what gets cited. These systems are opaque and change without notice treat them as working heuristics.
Which signal matters most?
Third-party corroboration appears to carry the most weight and is the hardest to build. Extractability produces the fastest gain per hour invested.
Can I improve all six at once?
Sequence them. Entity clarity first because it is a dependency, third-party corroboration started early because it is slowest.
How long before all six are working?
Six months for the fast signals to show and the slow ones to begin landing. Twelve for the compounding effect.
Does one strong signal compensate for a weak one?
Partly. Excellent comparative content with no third-party corroboration plateaus. The signals reinforce each other rather than substituting.
What if I have limited time?
Entity clarity, then extractability, then comparison pages. That is roughly 30 hours and covers the three cheapest signals with the fastest effect.