Answer in the opening lines
DecisiveModels retrieve a passage, not a page. If the answer is not in the first two sentences of a section, it usually is not in the chunk that gets quoted.
Every check we run, and why each one changes whether an answer engine cites you. We detect what a page actually is — its shape and its subject — then run the checklists that apply. A cyclone report and a lipstick page are not graded the same way.
Run against every page regardless of what it is.
Models retrieve a passage, not a page. If the answer is not in the first two sentences of a section, it usually is not in the chunk that gets quoted.
Googlebot renders JavaScript; GPTBot, ClaudeBot and PerplexityBot do not. JS-injected content is invisible to every answer engine at once.
Retrieval matches a user’s question against your headings. A heading that is already the question is the strongest match you can offer.
Controlled testing (the Princeton GEO study) found citing authoritative sources measurably increases how often a page is quoted. Models repeat attributed claims more readily than floating ones.
A number is self-contained and survives being lifted out of context. "Traffic fell sharply" is unusable to an answer engine; "traffic fell 58%" is quotable on its own.
A retrieved chunk arrives without the paragraph that defined "it". Sections that lean on pronouns become unusable once separated.
Answer engines prefer sources they can date, and heavily discount undated pages on anything time-sensitive.
Roughly two-thirds of the pages Google AI Mode cites, and about seven in ten of the pages ChatGPT cites, carry schema markup. It is the clearest signal in the whole checklist: it tells the engine what the page is instead of making it guess.
Engines retrieve a passage of roughly 100–300 words, not the page. One long undivided block gets split mid-thought and usually discarded; a page of tight, self-contained sections gives the engine something it can lift intact.
Pages carrying original figures — your own testing, survey or internal data — are measurably more likely to be quoted, because they are the only source for that number. Content that only restates what other sites already say gives an engine no reason to pick it over them.
The heading tree is how a parser works out which text belongs to which topic. Several H1s, or an H2 that jumps straight to H4, produce sections attached to the wrong heading — so the right answer gets filed under the wrong question.
The crawlers behind ChatGPT, Claude and Perplexity read HTML, not pictures. Anything that exists only inside an image — a spec table, a chart, a comparison — is simply absent unless the alt text says what it shows.
AI crawlers are far less patient than Googlebot — many give up after one to five seconds. A very heavy document risks being abandoned before the content is read, which costs the citation regardless of how good the writing is.
nosnippet and max-snippet:0 tell search engines not to show any text preview of this page at all — which also removes it as a candidate for AI Overviews and AI Mode. This is usually set by mistake, in a template, not per page.
AI crawlers do not watch video. Anything said only on camera — a demo, an explanation, a verdict — is invisible unless the surrounding text or a transcript restates it.
Added when the page is detected as this shape.
FAQ markup maps one question directly onto one quotable answer — the single most extractable structure a page can offer.
An identifiable author is a trust signal engines weigh when choosing between two otherwise-similar sources.
Added when the page is detected as this shape.
AI shopping answers quote price. If the price is injected by JavaScript, the engine reports the product without one — or skips it.
Product markup removes guesswork: the engine reads price, currency and availability as declared values instead of inferring them from prose.
A declared count persuades humans; the review text persuades models. If you claim 1,200 reviews and render three, three is the entire evidence base an engine has.
Shopping answers filter on availability. An engine that cannot read stock status usually drops the product from the comparison.
Product questions are asked by brand and model. A page titled only "Wireless Earbuds" cannot be matched to the query.
"Can I return it" is among the most-asked pre-purchase questions, and it is usually buried on a separate policy page the engine never associates with the product.
Specs are what comparison questions are answered from. In an image or a script-built table, they do not exist to the engine.
Added when the page is detected as this shape.
"Best X under Y" questions are answered from listing pages. ItemList tells the engine this is a ranked set rather than one long article.
A recommendation is only quotable if the engine can pair a specific product name with a specific price.
Added when the page is detected as this shape.
The question is "should I buy this". A verdict in the opening lines is the passage most likely to be lifted verbatim.
Pros/cons are pre-chunked, balanced, quotable statements — close to ideal retrieval units.
First-hand testing is what separates a review an engine will trust from marketing copy it will ignore.
Review markup lets an engine attribute the verdict to a named reviewer and a numeric score rather than parsing it out of prose.
Added when the page is detected as this shape.
Step-by-step answers are reproduced almost verbatim by assistants — but only when the steps are list markup rather than paragraphs.
This markup declares the steps, timings and materials explicitly, which is what voice and assistant answers read from.
An assistant relaying instructions needs to state what is required before step one, or the answer is incomplete and gets passed over.
Added when the page is detected as this subject.
News answers are assembled from the lead. A lead missing the when or where cannot be used to answer a factual question about the event.
On news queries, recency is close to decisive. Without machine-readable timestamps an engine cannot tell whether your version is the current one.
Engines strongly prefer the outlet that names the official, agency or document over one that writes "sources said".
A dateline is how a wire story declares where it was filed, which is exactly what location-scoped questions are matched against.
Added when the page is detected as this subject.
Health is the category where engines are most conservative. An uncredentialed health page is routinely passed over in favour of one with a named clinician.
A review line with a date is the clearest trust marker a health page can carry, and it is machine-readable.
Linking the study, WHO or CDC rather than another blog is what lets an engine treat a health claim as supported.
Its absence reads as a risk signal to an engine deciding whether to surface health guidance at all.
Absolute language ("cures", "guaranteed") is a suppression trigger for health content across every major engine.
Added when the page is detected as this subject.
"Does it contain X", "is it safe for sensitive skin" are answered from the INCI list. As an image it does not exist to a crawler.
Nearly every beauty question is conditional — "for oily skin", "for curly hair". Without the qualifier the page cannot match the question.
Application steps are a distinct, highly-asked question and a separate citation opportunity from the product description.
"Dermatologically tested" with no testing body reads as marketing. Named substantiation is what makes a claim repeatable by an engine.
Shade availability is a common question, and it is usually rendered as swatch images with no text alternative.
Added when the page is detected as this subject.
Shopping answers increasingly include delivery time and payment options; both are usually rendered by script or hidden behind a tab.
Added when the page is detected as this subject.
Entertainment questions are almost entirely entity lookups. Unnamed people cannot be matched to "who directed…".
"Where can I watch X" is the highest-volume question in this category and needs an explicit platform name and date.
This markup declares the title, cast and release date as data rather than leaving them to be parsed out of prose.
Added when the page is detected as this subject.
Assistants relay one step at a time. A step reading "repeat with the rest" is meaningless once separated from its neighbours.
"How long does it take" is asked of nearly every how-to, and is answerable only if you state it.
Added when the page is detected as this subject.
"X vs Y" is how buying questions are actually asked. A review naming no alternative cannot be retrieved for any comparison.
A price with no date goes stale silently, and an engine repeating a stale price is a reason to stop trusting the source.
First-hand expertise is the differentiator engines use to separate a real review from aggregated marketing copy.
Added when the page is detected as this subject.
"How much does X cost" and "what is the interest rate" are the single most-asked fintech questions, and the numbers are usually rendered by a calculator widget rather than existing as text.
"Who can apply" or "am I eligible" gates every financial-product decision. Without an explicit answer, an engine cannot tell a user whether the product even applies to them.
Financial claims carry more scrutiny than most content. A named regulator or license number is what separates a credible source from marketing copy, for both readers and answer engines.
A page that only lists upside reads as an advertisement. Named, specific risk disclosure is a trust signal and is often a legal requirement.
Added when the page is detected as this subject.
"What does it cost" and "what is the rate per square foot" are the two numbers every property question needs, and listings often bury them in a downloadable brochure.
"Is this in X locality", "what configuration is available" and "when can I move in" are asked of every property page. Vague location or a missing possession date makes the listing unusable for these queries.
In markets with a property regulator, a listed registration number is the clearest trust signal a listing can carry — and its absence is itself a question buyers ask.
"What is nearby" — schools, hospitals, transit — is a standard property question, and is only answerable if named rather than shown on an interactive map widget alone.
Added when the page is detected as this subject.
"What engine does it have", "what is the mileage" — these are answered from a spec sheet. If it renders as an image or a JS-built table, the numbers do not exist to a crawler.
"Which variant is this" and "what does the on-road price come to" cannot be answered from an ex-showroom figure alone — a very common gap.
"Is it safe", "how many airbags" are pre-purchase questions that specifically need a named rating or feature list — general "safety" copy does not answer them.
Automotive buying decisions are almost always comparative. A page naming no competing model cannot be retrieved for any "X vs Y" query.
Added when the page is detected as this subject.
"Can I apply", "what does it cost" and "how long does it take" are asked before anything else about a course, and are frequently split across a brochure PDF instead of the page itself.
Course quality is judged by who teaches it. An unnamed "expert faculty" claim carries no evidence an engine can repeat.
"What will I be able to do after this" and "what are the outcomes" are the actual purchase questions — a syllabus list alone does not answer them.
Course markup lets an engine read subject, level, instructor and duration as structured data instead of parsing marketing copy.
Added when the page is detected as this subject.
"What does this tool do" and "who is it for" are the first two questions any comparison or recommendation prompt needs answered — and are often buried under a hero tagline that says neither.
"How much does it cost" is asked constantly in tool-comparison prompts, and pricing pages frequently render the actual numbers via JavaScript widgets that a crawler never sees.
"X vs Y" and "alternatives to X" are extremely common SaaS-discovery prompts. A product page that never names a competitor cannot be surfaced for either.
B2B buying prompts frequently ask about compliance (SOC 2, GDPR). Its absence from the page is itself an answer an engine will report.
Added when the page is detected as this subject.
"Who sings this", "what album is it from" and "when did it release" are the core entity questions for any music page, and are the first things an engine needs to ground an answer.
"Where can I listen to this" is a direct-action question. Naming the platforms (Spotify, Apple Music) is what makes the page useful to point to, rather than just descriptive.
MusicRecording/MusicAlbum markup declares artist, album and duration as structured data an engine can read directly.
Run one against this checklist. The first check is on us.
Run a check