Every check we run
156 checks, every one of them wired into the same scan you'd run on your own page — 123 classic SEO checks and 33 AI-search-readiness checks. This list is generated straight from our own check registry, so it can't drift from what actually ships.
Curious how these roll up into your score? See our methodology . Want to see them run for real? Run a free check .
156
Total checks
123
SEO checks
33
AI-readiness (GEO) checks
Classic SEO — 123 checks
Content Quality
Whether a page says enough, says it clearly, and says something a reader could not get from a thinner competitor — thin content, readability, duplication, and filler language.
10
Content Quality
Whether a page says enough, says it clearly, and says something a reader could not get from a thinner competitor — thin content, readability, duplication, and filler language.
- critical
Content couldn't be analyzed
The scanner couldn't parse the page's content at all, which usually means the URL isn't returning real, crawlable HTML.
- critical
Thin content
Pages with too few words rarely cover a topic in enough depth to rank well or fully answer what a visitor came looking for.
- high
Low readability
Long, complex sentences and heavy jargon make content harder to scan and understand, for both readers and ranking systems.
- medium
AI-generic filler language
Generic filler phrases like "let's dive in" or "unlock the power of" read as low-effort AI writing and dilute concrete, useful content.
- high
Duplicate content risk
Near-duplicate pages compete against each other for the same rankings and dilute the signals search engines use to pick a winner.
- medium
Paywalled or gated content
Content locked behind a subscribe or sign-in wall can go unindexed unless flexible-sampling structured data tells search engines what is behind it.
- medium
H1 doesn't match page content
A heading that doesn't reflect what the body actually discusses is a topical-relevance red flag for both readers and search engines.
- low
Low text-to-HTML ratio
A page that's mostly markup with little visible text gives search engines very little substantive content to evaluate.
- medium
No specific claims or data
Concrete numbers, dates, and sourced facts read as more trustworthy and citable than vague, generic statements.
- low
No first-hand experience signals
Language showing real testing or direct experience ("I tested", "we found") signals the content reflects genuine use, not just aggregated research.
Technical SEO
The foundation search engines need before content quality even matters: crawlability, indexability, HTTPS, redirects, robots directives, and security headers.
27
Technical SEO
The foundation search engines need before content quality even matters: crawlability, indexability, HTTPS, redirects, robots directives, and security headers.
- critical
Page returns an HTTP error
A page returning a 4xx or 5xx status cannot be indexed at all — it needs to load successfully or redirect properly to the right destination.
- high
Unresolved redirect
Internal links and canonical references pointing through a redirect hop instead of the final URL waste crawl budget and dilute link equity.
- critical
Not served over HTTPS
HTTPS is a confirmed Google ranking signal, and browsers actively warn visitors away from HTTP pages, hurting trust and conversions.
- critical
Noindex directive present
A noindex tag or header removes the page from search results entirely, even if everything else about it is fine.
- critical
robots.txt blocks all crawlers
A site-wide Disallow rule stops search engines from crawling any page at all, silently taking the whole site out of search.
- critical
Page blocked by robots.txt
A page-specific Disallow rule can block Googlebot from an important page just as completely as a site-wide block, and is often shipped by accident.
- high
robots.txt returns a server error
When robots.txt itself fails, Google halts crawling of the whole site for hours and falls back to a stale cached copy for up to 30 days.
- medium
Slow server response time
A slow time-to-first-byte delays everything downstream — page rendering, Core Web Vitals, and how much of the site crawlers can get through.
- high
Soft 404 page
A missing page that returns HTTP 200 instead of a real 404 wastes crawl budget and confuses search engines about what actually exists.
- medium
Inconsistent www/non-www redirect
When both domain variants don't consistently redirect to one canonical version, search engines can split ranking signals between the two.
- low
Inconsistent trailing-slash handling
Serving both a trailing-slash and non-trailing-slash URL as separate live pages creates avoidable duplicate-content confusion.
- high
Missing security headers
Without headers like CSP, HSTS, and X-Frame-Options, the site is more exposed to clickjacking, downgrade attacks, and injected content.
- medium
Mixed content on HTTPS page
Resources loaded over plain HTTP on an HTTPS page trigger browser warnings and can be silently blocked, breaking the page for visitors.
- high
Form submits over insecure HTTP
Submitting form data — including credentials or personal details — to a plain http:// endpoint exposes it to interception in transit.
- medium
Missing XML sitemap
Without a sitemap.xml, search engines have to rely purely on link discovery to find and prioritize the pages worth crawling.
- medium
Malformed sitemap
A sitemap that isn't well-formed XML is typically ignored entirely by search engines, providing no crawl-discovery benefit at all.
- medium
Requires JavaScript to render
Some crawlers and link-preview bots don't execute JavaScript, so they see a near-empty page unless it's rendered server-side.
- critical
Leaked credential in page source
A live API key or credential visible in HTML or JS is publicly exposed to anyone who views source and should be rotated immediately.
- medium
Hreflang tag issue
Inconsistent or missing hreflang declarations can send search engines the wrong language or region variant for a given searcher.
- high
Uncrawlable links
Links built with javascript: hrefs, #-only anchors, or empty hrefs can't be followed by crawlers, hiding whatever they point to.
- high
Risky redirect signal
Cross-domain redirects and instant meta-refreshes read as manipulative signals rather than a legitimate same-site 301 redirect.
- medium
Cross-domain canonical URL
A canonical pointing at a different domain tells search engines the authoritative version of this content lives elsewhere entirely.
- low
No analytics tool detected
Without an analytics tool installed, there's no way to measure traffic, conversions, or search performance for this page.
- medium
Unsafe target="_blank" links
A target="_blank" link without rel="noopener" lets the destination page access window.opener and redirect the original tab (reverse tabnabbing).
- low
Exposed mailto email address
Plain-text mailto links are easy targets for automated spam-harvesting bots scraping email addresses off the page.
- high
Dead-end page with no links
A page with zero internal or external links gives users nowhere to go next and crawlers no way to discover more of the site from it.
- medium
Links point to localhost
A leftover dev or staging link to localhost is broken for every real visitor and crawler outside the developer's own machine.
On-Page SEO
The classic ranking signals on the page itself — titles, meta descriptions, headings, canonical tags, and Open Graph data.
22
On-Page SEO
The classic ranking signals on the page itself — titles, meta descriptions, headings, canonical tags, and Open Graph data.
- critical
Meta tags couldn't be analyzed
The page couldn't be parsed for meta tags at all, which usually means the URL isn't returning real, crawlable HTML.
- critical
Missing title tag
The title tag is one of the strongest on-page ranking and click-through signals — without one, the page has no clear SERP headline.
- medium
Title tag too short
A too-short title wastes valuable SERP real estate and often fails to include enough context or the primary keyword to attract clicks.
- low
Title tag too long
Titles over roughly 60 characters get truncated in search results, potentially cutting off the most compelling or relevant part.
- high
Missing meta description
Without a meta description, search engines generate their own snippet from page text, which is often less compelling than a written one.
- low
Meta description too long
A meta description over about 160 characters gets truncated in search results, cutting off part of the summary before it can persuade a click.
- medium
Missing canonical tag
Without a canonical URL, search engines have to guess which version of a page is authoritative, risking duplicate-content dilution.
- low
Generic anchor text
Link text like "click here" or "read more" tells users and search engines nothing about the destination, unlike descriptive anchor text.
- low
Outdated meta keywords tag
Google has ignored the meta keywords tag for ranking since 2009, and its presence reads as an outdated or spam-adjacent SEO approach.
- critical
Missing H1 heading
The H1 is the clearest on-page signal of a page's main topic to both search engines and readers scanning the page.
- medium
Multiple H1 headings
More than one H1 muddies the single-topic hierarchy a page should present, making the primary subject less clear.
- low
Deprecated HTML tag
Tags like <center> or <font> have been obsolete in HTML for years and should be replaced with modern CSS equivalents.
- high
Heading structure couldn't be checked
Without parseable HTML, H1 presence and heading hierarchy can't be verified at all, hiding a potentially serious structural issue.
- low
Missing Twitter Card tags
Without twitter:card meta tags, links to this page won't preview correctly when shared on X/Twitter, hurting social click-through.
- low
No social profile links
Linking to official social profiles supports Organization schema's sameAs properties and general brand-credibility signals.
- medium
Non-descriptive URL path
A URL built from numeric IDs or query strings gives users and search engines no readable clue about the page's content before clicking.
- medium
Page type mismatch with SERP
When a page's format doesn't match what's currently ranking for its target keyword, it's competing against a content type Google prefers there.
- medium
Missing html lang attribute
Without a lang attribute, accessibility tools, translation prompts, and language-targeted search results can't reliably identify the page's language.
- medium
Duplicate canonical tags
Multiple conflicting canonical tags confuse search engines about which URL is actually authoritative for this content.
- low
Missing og:title tag
Without an og:title tag, the page has no clear title when shared on social media, hurting how links appear when posted.
- low
Missing og:description tag
Without an og:description, shared links fall back to arbitrary page text instead of a meaningful, chosen description.
- medium
Missing og:image tag
Links shared without an Open Graph image get far less engagement than ones with a representative preview image.
Structured Data
Whether your JSON-LD schema markup is present, valid, and current — deprecated or malformed schema can silently lose rich results.
13
Structured Data
Whether your JSON-LD schema markup is present, valid, and current — deprecated or malformed schema can silently lose rich results.
- low
Structured data couldn't be checked
Without parseable HTML, structured data markup can't be verified at all, hiding any schema issues that might exist.
- high
No structured data found
Without JSON-LD markup, the page misses out on rich results and gives search engines no explicit signal about what type of content it is.
- high
Deprecated schema type
A no-longer-supported schema type provides no ranking or rich-result benefit and should be removed or replaced.
- medium
SERP-retired schema type
Google no longer generates a rich result for this schema type, so it can't be relied on for search-result appearance anymore.
- medium
Schema uses HTTP @context
A JSON-LD block declaring an http:// @context instead of https:// is an easy, low-effort schema-hygiene fix.
- medium
Missing required schema field
A schema block missing a required field disqualifies the page from that rich-result type entirely until the field is added.
- medium
Invalid breadcrumb markup
BreadcrumbList markup needs sequential positions, names, and item URLs to actually qualify for the breadcrumb rich result.
- medium
Low-quality FAQ markup
FAQPage schema needs a real name and a substantive, non-duplicate answer for every question to be useful for rich-result or AI citation.
- low
Client-side rendering blind spot
If schema markup is only injected via JavaScript, it may never reach the served HTML that non-JS-executing crawlers actually see.
- medium
Article headline too long
Google disqualifies an Article schema block from its rich result entirely once the headline exceeds roughly 110 characters — it isn't wrapped, just dropped.
- medium
Article schema missing image
Google requires an image of at least 1200x800px for Article-family rich results, so a missing image blocks eligibility outright.
- medium
Article schema image too small
An Article schema image below Google's minimum dimensions fails the rich-result eligibility bar even though an image is present.
- medium
Article schema violation
The Article schema block fails one of Google's structural requirements, blocking it from rich-result eligibility until adjusted.
Core Web Vitals
Real-user loading and interactivity metrics — LCP, CLS, FCP, TTFB, and INP — the speed signals that affect both ranking and conversion.
6
Core Web Vitals
Real-user loading and interactivity metrics — LCP, CLS, FCP, TTFB, and INP — the speed signals that affect both ranking and conversion.
- low
Core Web Vitals not measured
Until a real Core Web Vitals measurement runs, there is no data on how the page actually performs for real users on LCP, CLS, or INP.
- critical
Poor Largest Contentful Paint
A slow LCP means visitors wait too long to see the page's main content load, which is both a UX and a Google ranking signal.
- critical
Poor Cumulative Layout Shift
Content that jumps around as the page loads frustrates users and can cause accidental clicks on the wrong element.
- medium
Poor First Contentful Paint
A slow first paint leaves visitors staring at a blank screen longer than necessary before anything renders at all.
- medium
Poor Time to First Byte
A slow server response delays every subsequent rendering milestone, dragging down the whole page-load experience.
- medium
Poor Interaction to Next Paint
Slow response to clicks and taps makes a page feel sluggish and unresponsive even after it has visually finished loading.
AI Search Readiness
Whether a page is structured so AI assistants can find, parse, and safely cite it — frontloaded answers, entity density, freshness signals, and llms.txt.
10
AI Search Readiness
Whether a page is structured so AI assistants can find, parse, and safely cite it — frontloaded answers, entity density, freshness signals, and llms.txt.
- medium
No front-loaded answer
AI engines prefer a direct answer in the first 40-60 words of a section — content that builds up to its point is less likely to be lifted and cited.
- medium
Low entity density
Named entities like people, places, organizations, and brands are what AI systems use to ground and cite content — generic terms give them nothing to anchor to.
- low
No statistics detected
Numeric evidence — specific dates, counts, percentages, or measured results — improves how likely AI systems are to cite a passage as reliable.
- medium
No freshness signal
Without a visible published/updated date and matching structured data, AI search systems have no recency signal to score the content on.
- low
No author signal
AI systems weigh clear authorship as a citation-worthiness signal, so content with no visible byline is less likely to be trusted and cited.
- medium
AI readiness signals unavailable
Without parseable HTML, AI search readiness signals like entity density and answer-first structure can't be assessed at all.
- medium
Missing llms.txt file
An llms.txt file is a low-effort, standardized way to guide AI crawlers to a site's most important content per the llmstxt.org spec.
- low
llms.txt missing site name
Without a valid H1 site name as the first line, AI crawlers reading llms.txt may not reliably identify which site it belongs to.
- high
AI crawlers blocked in robots.txt
Blocking AI crawlers like GPTBot or ClaudeBot in robots.txt prevents this content from ever being cited by AI assistants at all.
- low
Site files unavailable
Without a successful domain-level fetch, llms.txt and robots.txt AI-crawler access can't be checked for this scan.
Images
Alt text, image presence, and the accessibility and SEO signal a page loses when images are missing or unlabelled.
7
Images
Alt text, image presence, and the accessibility and SEO signal a page loses when images are missing or unlabelled.
- low
Images couldn't be checked
Without parseable HTML, image alt text, formats, and loading attributes can't be audited at all.
- low
No images on page
Relevant visuals — screenshots, diagrams, product photos — can improve engagement and give AI and image search another way to reference the page.
- high
Missing image alt text
Alt text helps screen readers describe images to visually impaired visitors and can drive additional traffic through image search.
- medium
Legacy image format
Older formats like JPEG and PNG are typically 25-50% larger than WebP or AVIF at equivalent quality, slowing down page load.
- critical
Lazy-loaded hero image
Lazy-loading the above-the-fold hero image delays exactly the element Largest Contentful Paint measures, directly hurting that Core Web Vital.
- medium
Missing responsive image srcset
Without a srcset attribute, browsers always load the largest image version regardless of viewport size, wasting bandwidth and slowing load.
- medium
Video issue detected
Problems with embedded video, such as missing captions or a broken embed, can limit both accessibility and citability of that content.
Authority & Trust Signals
E-E-A-T signal strength and internal linking health — orphan pages and thin internal linking quietly undermine authority even when the content itself is fine.
5
Authority & Trust Signals
E-E-A-T signal strength and internal linking health — orphan pages and thin internal linking quietly undermine authority even when the content itself is fine.
- low
Trust signals unavailable
Without parseable HTML, spam and trust checks like cloaking, hidden text, and E-E-A-T signals can't be evaluated at all.
- medium
Orphan page risk
A page with no internal links pointing to it, and that isn't the homepage, is hard for both crawlers and users to discover through normal navigation.
- low
Few internal links
A page with only a handful of internal links pointing to it has weaker internal-linking signals than better-connected pages on the site.
- critical
Fails YMYL E-E-A-T standards
For Your-Money-or-Your-Life topics like health or finance, content without a named expert author or that reads as unreviewed AI output can be effectively disqualified from ranking.
- high
Weak E-E-A-T signals
Weak trustworthiness, expertise, authoritativeness, or experience signals in the content make it a harder sell for search engines to rank highly, especially on sensitive topics.
Spam & Local Business Signals
Two different risks in one bucket: abuse patterns that can trigger a manual action (cloaking, hidden text, keyword stuffing, link spam, scaled content), and — where relevant — local-business completeness like NAP consistency, hours, and reviews.
23
Spam & Local Business Signals
Two different risks in one bucket: abuse patterns that can trigger a manual action (cloaking, hidden text, keyword stuffing, link spam, scaled content), and — where relevant — local-business completeness like NAP consistency, hours, and reviews.
- critical
High cloaking risk
Serving different content to Googlebot than to real users is a manual-action risk that can get a site penalized outright.
- medium
Possible cloaking signal
Page behavior that differs by user agent is worth reviewing, since unintentional cloaking still carries the same penalty risk as deliberate cloaking.
- high
Hidden text detected
Hiding keyword-stuffed content from users while showing it to crawlers is a manual-action risk under Google's spam policies.
- medium
Keyword stuffing suspected
Repetitive keyword or location lists read as manipulative to both readers and Google's spam detectors, rather than as genuine, useful content.
- high
High doorway page risk
Thin content with high link density and geographic title variation is a classic doorway-page pattern that Google's spam policies specifically target.
- medium
Some doorway page signals
Thin content, a canonical pointing elsewhere, or unusual title variation are early signals of a page existing mainly to funnel visitors elsewhere.
- high
High link spam risk
An excessive external-link ratio or unmarked affiliate links violate Google's link-scheme guidelines and put the page at manual-action risk.
- medium
Some link spam signals
Elevated footer or sidebar link density and unmarked affiliate links are worth reviewing before they escalate into a full link-spam risk.
- medium
UGC links missing rel attributes
Links in comments or forums need rel="ugc" or rel="nofollow" so search engines don't attribute user-submitted links' authority to the site.
- high
Thin affiliate content
Google's guidelines require affiliate pages to add real value beyond just linking to merchants — too little original content around multiple affiliate links risks a ranking penalty.
- medium
Affiliate links without review language
Affiliate links without first-hand-experience or review language read as a plain link list rather than a genuine, helpful review.
- high
Scaled/low-value content risk
Low readability, repetitive paragraphs, and no named entities are the fingerprint of mass-produced content that Google's spam policies specifically target.
- medium
Some scaled-content signals
Low readability or a lack of named entities can make content read as mass-produced even before it crosses into full scaled-content risk.
- critical
Malware signals detected
Obfuscated scripts are a strong indicator of a site compromise and need immediate investigation before anything else about the page matters.
- medium
Suspicious third-party iframes
An iframe embedding content from an unrecognized or untrusted domain is a common vector for malware and compromised-site injections.
- low
Third-party author byline
An author byline linking to a different domain than the one being scanned needs confirming as a genuine syndication arrangement, not a reputation-abuse pattern.
- medium
Undisclosed sponsored content
Sponsored or partnership language without a corresponding rel="sponsored" link fails to properly disclose paid content to both readers and search engines.
- medium
Low E-E-A-T signal count
Missing trust markers like an author byline, About/Contact links, or Organization schema are especially costly on YMYL topics where trust signals carry real ranking weight.
- medium
Incomplete NAP data
Incomplete Name/Address/Phone data in local business schema hurts local-pack ranking eligibility, even when the schema type itself is present.
- low
Missing opening hours
Without an openingHours property, search engines can't show accurate business hours in local search results.
- low
Missing geo coordinates
Local business schema without latitude/longitude coordinates reduces eligibility for map and local-pack placements.
- medium
References a removed GBP feature
Mentioning a Google Business Profile feature that has since been removed misleads readers about capabilities that no longer exist.
- low
Low review count
Review count and recency both factor into local ranking and review-rich-result eligibility, and this page falls below the informal threshold some platforms use.
AI Search Readiness (GEO) — 33 checks
Whether ChatGPT, Claude, Gemini, and Perplexity can actually find, parse, and cite your page — the genuinely differentiated part of what we check.
Brand Authority Signals
Whether an AI system can tell who wrote this and trust the brand behind it — author attribution, entity consistency, named entities, and organization schema.
5
Brand Authority Signals
Whether an AI system can tell who wrote this and trust the brand behind it — author attribution, entity consistency, named entities, and organization schema.
- medium
Missing author/organization attribution
A visible author or organization byline, mirrored in structured data, is what lets AI systems attribute content to a named, credible source.
- medium
Inconsistent brand/entity naming
When a brand name doesn't appear consistently across title, meta description, and H1, AI systems have a harder time reliably associating the page with that brand.
- high
Missing named entities
Naming specific people, companies, products, or places is what AI systems use to ground and cite content, rather than generic terms.
- medium
Missing organization schema
Organization (or Brand/LocalBusiness) structured data gives AI systems and search engines a machine-readable identity to attribute content to.
- medium
Missing statistics or numeric claims
Specific, verifiable numbers make claims easier for AI systems to trust and cite compared to vague, unsupported statements.
Citability
The passage-level traits that make a paragraph quotable by an AI answer — answer-first structure, evidence-backed claims, ideal length, freshness, and low hedging.
10
Citability
The passage-level traits that make a paragraph quotable by an AI answer — answer-first structure, evidence-backed claims, ideal length, freshness, and low hedging.
- high
No early, front-loaded answer
Rewriting key sections to lead with the direct answer, rather than building up to it, makes content far more likely to be lifted and cited by AI systems.
- medium
Stale content
Content untouched for 180+ days, and especially a year or more, reads as increasingly stale to AI systems weighing recency in what to cite.
- medium
Unsupported superlative claims
Claim words like "best", "proven", or "guaranteed" without a specific number, stat, or citation nearby are far less likely to be cited by AI answer engines.
- medium
No freshness signal
A visible published/updated date, matched to datePublished/dateModified in structured data, is what AI search systems use for recency scoring.
- medium
Answer not front-loaded / off sweet spot
AI citations concentrate heavily in the first 30% of a page, and total length in the roughly 800-1500 word range has the strongest empirical AI-extraction coverage.
- high
Passages not sized for citation
Passages in the 134-167 word range are long enough to carry context but short enough for an AI system to quote cleanly as a self-contained answer.
- low
High hedge-word density
Hedging language like "might", "could", or "it seems" makes claims less likely to be lifted as a confident, quotable statement by an AI system.
- medium
Stale statistics or outdated tech
Citing dated studies as current, or referencing long-obsolete technology, undermines the credibility of a passage an AI system might otherwise cite.
- medium
Passages too long for citation
Long, sprawling paragraphs are harder for an AI system to extract and quote cleanly than passages broken into a citable, self-contained length.
- medium
Low statistical density
Content dense in concrete, sourced facts and certifications is measurably more likely to be cited by AI answer engines than purely descriptive text.
Multi-Modal Readiness
Whether a page gives an AI system more than plain text to work with — data tables, images, and video content it can also draw on.
3
Multi-Modal Readiness
Whether a page gives an AI system more than plain text to work with — data tables, images, and video content it can also draw on.
- low
No data table
A structured data table presents tabular information — pricing, specs, comparisons — in a form that is easy for both readers and AI systems to extract.
- low
Too few images
Relevant visuals like diagrams, screenshots, or photos give AI and image search another way to reference and cite the page.
- low
No video content
A relevant video, where it genuinely fits the content, is a strong multi-modal signal that broadens how the page can be discovered and cited.
Structural Readability
Semantic HTML that helps an AI system parse a page correctly — proper landmarks, lists, sections, and genuine question-and-answer structure.
6
Structural Readability
Semantic HTML that helps an AI system parse a page correctly — proper landmarks, lists, sections, and genuine question-and-answer structure.
- low
Missing semantic <article> element
Wrapping core content in a semantic <article> element gives crawlers and AI systems a clear structural signal about where the primary content lives.
- low
No lists for scannability
Bulleted or numbered lists break up scannable information like steps, features, or comparisons in a form both readers and AI systems parse easily.
- medium
Missing semantic <main> landmark
A semantic <main> element clearly marks the primary page content, helping both accessibility tools and content-extraction systems locate it.
- medium
Question headings not directly answered
Question-style headings whose following text doesn't directly and immediately answer the question are less likely to be lifted cleanly by an AI system.
- medium
No question-per-section structure
Framing key topics as a question subheading with the answer directly beneath it matches exactly the shape AI answer engines extract from.
- medium
Missing semantic <section> elements
<section> elements dividing distinct topics give crawlers and AI systems a clearer structural map of the content than undifferentiated markup.
AI Crawler Access
Whether AI crawlers and citation bots can actually reach and read a page at all — bot access, llms.txt, markdown negotiation, and content-signal headers.
9
AI Crawler Access
Whether AI crawlers and citation bots can actually reach and read a page at all — bot access, llms.txt, markdown negotiation, and content-signal headers.
- high
AI crawlers blocked in robots.txt
Blocking AI crawlers in robots.txt directly prevents this content from being citable by whichever AI systems are disallowed.
- medium
No AI-crawler content negotiation
Recognizing an AI-crawler user agent and serving it a lighter, stripped-down response is an emerging competitive edge for AI citation, though most sites don't do this yet.
- high
Citation-intent AI bots blocked
Blocking bots that answer live user queries right now, like PerplexityBot or ChatGPT-User, has a high, immediate impact on real-time AI-answer-engine visibility.
- medium
Content-Signal opts out of AI retrieval
A robots.txt Content-Signal directive of ai-input=no explicitly opts out of the real-time AI retrieval mechanism that AI-Overview-style answers pull live content through.
- medium
Missing llms.txt file
An llms.txt file is a low-effort, low-risk addition per the llmstxt.org spec, though recent log studies found most llms.txt files receive essentially zero AI-crawler requests.
- low
No Accept: text/markdown negotiation
Honoring an Accept: text/markdown request header lets AI crawlers request clean content directly instead of parsing it out of full HTML.
- low
No AI-markdown-twin route
Serving a plain-Markdown twin of a page at {path}.md is a cheap, high-signal convention that lets AI agents fetch clean content without HTML stripping.
- low
Missing RSL licensing file
A /.well-known/rsl.json file declares content-licensing terms for AI training — an emerging, low-effort signal for sites with a clear licensing stance.
- high
Training-intent AI bots blocked
Blocking model-training crawlers has no immediate citation impact and may be a deliberate content-rights choice rather than something that needs fixing.