A page can look perfect in a browser and still be invisible to Google. One line in robots.txt, a tag left over from the staging site or a redirect loop is enough, and none of them shows when you click around. This technical SEO checklist is built to catch them.
Each check comes with the free tool to test it and what a pass looks like, based on Google Search Central’s documentation. It’s the technical layer of our guide to SEO-friendly websites.
The technical SEO checklist at a glance
| Question | What to check | Where to check it | What a pass looks like |
|---|---|---|---|
| Can it be crawled? | robots.txt, sitemap, internal links | robots.txt report, Sitemaps report, a crawler | Key pages allowed, linked and listed |
| Can it be indexed? | noindex rules, status codes | URL Inspection, Page indexing report | Key pages indexed; every exclusion deliberate |
| One version of each page? | HTTPS, www, canonicals, parameters | Address bar, crawler, URL Inspection | One URL per page, and Google agrees |
| Mobile and rendering? | Mobile content, JavaScript | Lighthouse, URL Inspection live test | Full content on mobile and in the rendered HTML |
| Redirects and 404s? | Chains, soft 404s, broken links | Crawler, Page indexing report | One hop per redirect; no links to missing pages |
Three free tools do most of the work: Google Search Console, set up as a Domain property (verified through your DNS) so it covers every subdomain and both http and https; a desktop crawler such as Screaming Frog SEO Spider, whose free version crawls up to 500 URLs; and your browser’s view source and developer tools. Report names are correct at the time of writing (September 2026).
Can search engines crawl your site?
robots.txt
The robots.txt file at the root of your domain tells crawlers which paths not to fetch. On a typical WordPress site, a healthy one is short:
User-agent: *
Disallow: /wp-admin/
Allow: /wp-admin/admin-ajax.php
Sitemap: https://www.example.com/sitemap.xml
The most damaging version is Disallow: / under User-agent: *, a staging leftover that blocks everything. And robots.txt stops crawling, not indexing: a blocked URL can still appear in results, without a description.
Check it. Open yourdomain.com/robots.txt, then the robots.txt report in Search Console’s settings, which shows the version Google last fetched and any errors.
What good looks like. Only areas with no search value are blocked, such as admin screens and internal search results, and never the CSS or JavaScript a page needs to render. The file loads reliably: Google treats a missing robots.txt (a 404) as “no restrictions”, but a server error can make it stop crawling the whole site until the file loads again.
XML sitemap
A sitemap helps Google discover pages but doesn’t force anything into the index.
Check it. Submit it in Search Console’s Sitemaps report, confirm the status reads Success, then crawl its URLs as a list.
What good looks like.
- Every URL returns 200, is indexable and is the canonical version.
- The count roughly matches the pages you want found. A 40-page website with 900 sitemap URLs is probably listing tag archives or image attachment pages.
lastmodchanges only when the content does. Google says it ignorespriorityandchangefreq, and useslastmodonly when it’s consistently and verifiably accurate.
Internal links
Google follows links written as <a> elements with an href address. A “link” that only works through a JavaScript click handler may never be followed.
Check it. Crawl from the home page and compare the results with your sitemap: sitemap URLs the crawl never reached are orphan pages.
What good looks like. Key service pages sit a few clicks from the home page (check the crawl depth column), linked with descriptive anchor text. Internal linking and site structure explains how to plan it.
Can your pages be indexed?
noindex rules and status codes
A noindex rule sits in the page as <meta name="robots" content="noindex">, or in an X-Robots-Tag HTTP header that never shows in the page source. The usual accidents are a staging setting carried over at launch, WordPress’s “Discourage search engines from indexing this site” box (Settings → Reading) left ticked, or an SEO plugin hiding a whole content type. Google has to crawl a page to see its noindex, so don’t also block it in robots.txt.
Status codes must tell the truth: 200 for live pages, 301 (or 308) for permanent moves, 404 or 410 for pages that are gone. Watch for soft 404s (pages that say “not found” but return 200) and repeated server errors, which make Google slow its crawling.
Check it. URL Inspection in Search Console shows “Indexing allowed?” for one URL; your crawler’s indexability and status code columns cover the whole site.
What good looks like. Live pages return 200, and noindex appears only on pages that don’t belong in search, such as thank-you and login pages.
The Page indexing report
This report counts the pages Google knows on your site and, for those it hasn’t indexed, gives a reason with example URLs. The skill is knowing which reasons are normal.
- Usually fine: “Page with redirect”, “Alternate page with proper canonical tag”, “Not found (404)” for pages you removed, and noindex exclusions on pages you meant to hide.
- Worth investigating: robots.txt blocks on pages you want found, “Duplicate, Google chose different canonical than user”, “Soft 404”, “Server error (5xx)” and “Redirect error”.
- A judgement call: “Crawled – currently not indexed” (fetched but not indexed, with no reason given: ask whether the page adds anything your other pages don’t) and “Discovered – currently not indexed” (known but not yet crawled).
What good looks like. Every important page is indexed, and you can explain every exclusion. If a key page is missing, why a website isn’t showing on Google works through the diagnosis.
Is there one version of each page?
To a search engine, http://example.com/services, https://www.example.com/services/ and the same address with ?utm_source=newsletter can be three separate pages. Duplicates split signals such as links and leave Google to pick one.
HTTPS and www
Check it. Type each variant of one inner page into the address bar (with and without www, http or https, trailing slash or not), then crawl that list to see the codes and hops.
What good looks like. Every variant reaches one preferred version through a single 301. Internal links, canonical tags and the sitemap all use that version, and the browser console shows no mixed-content warnings.
Canonical tags
A canonical tag names a page’s preferred URL. Google treats it as a strong hint rather than an order, so each indexable page should point to itself and every other signal should agree. Mistakes to hunt for:
- A template bug that points every canonical at the home page.
- Canonicals pointing to URLs that redirect, return 404, carry noindex or still use the staging domain.
- Paginated pages canonicalised to page 1. Google advises giving each page in a series its own canonical.
- Translations canonicalised to the main language, instead of pointing to themselves and linking to each other with hreflang (see the multilingual SEO guide).
Check it. URL Inspection shows the “User-declared canonical” and “Google-selected canonical”. On every important page, they should match.
URL parameters
Tracking tags, sorting and filters (?sort=price) create new URLs for the same content. Google retired Search Console’s URL Parameters tool in 2022, so the fix lives on the site: parameter versions should canonicalise to the clean URL, and internal links should never carry tracking tags, which also muddle your analytics.
Does it work on a phone and render properly?
Mobile-first indexing
Google crawls and indexes with its smartphone crawler, so the mobile version of each page is what counts. Responsive sites serve the same HTML to every screen; the risk lies in separate mobile templates and designs that drop sections, headings or structured data on small screens. Googlebot doesn’t click, tap or swipe, so content that only loads after an interaction may be missed; text already in the HTML but inside tabs or accordions is fine.
Check it. Google retired its Mobile-Friendly Test in December 2023. Use Lighthouse in Chrome’s developer tools, your own phone on mobile data, and URL Inspection’s live test, which screenshots the page as Googlebot renders it.
What good looks like. Nothing you want ranked exists only on desktop.
JavaScript rendering
Google renders pages in a recent version of Chrome, though not always straight after crawling, and not every crawler runs JavaScript. Pages delivered as complete HTML, as standard WordPress pages are, avoid most of the risk.
Check it.
- Open View source, not Inspect. If your headings, text and menu links are there, the rendering risk is low.
- If not, run URL Inspection’s live test and open View tested page: the HTML tab shows what Google rendered, and “More info” lists resources it couldn’t load.
- Without Search Console access, Google’s Rich Results Test shows the rendered HTML for any public URL.
What good looks like. The rendered HTML contains the full text, real <a href> links and the correct title, canonical and robots tags, and each page has its own URL rather than a # fragment.
Speed and Core Web Vitals
Speed deserves its own project, but know the bar: Google rates a page good at a Largest Contentful Paint of 2.5 seconds or less, Interaction to Next Paint of 200 milliseconds or less and Cumulative Layout Shift of 0.1 or less, at the 75th percentile of real visits. Core Web Vitals explained covers each metric, and the performance guides go further.
Are redirects and 404s under control?
Redirect chains and loops
A chain is A redirecting to B, which redirects to C. Chains build up quietly through HTTPS moves, restructures and renamed pages, and each hop slows visitors. Googlebot follows up to 10 hops, but Google advises redirecting straight to the final page.
Check it. Your crawler’s redirect report lists chains and loops; ones Google gave up on appear in the Page indexing report as “Redirect error”.
What good looks like. Every redirect reaches its final page in one hop, and internal links point straight to final URLs. For redesigns and domain changes, the website migration SEO checklist covers redirect mapping.
Removed pages and broken links
Redirect a removed page to its closest equivalent, or let it return 404 or 410 if there isn’t one. Sending every missing URL to the home page feels tidy, but it confuses visitors, and Google says it may treat those redirects as soft 404s.
Check it. Filter your crawl for 4xx responses and see which pages link to them.
What good looks like. No internal links to missing pages, and a helpful custom 404 page that returns a real 404 status.
What to fix first
- Blockers: fix today. A site-wide noindex or
Disallow: /, robots.txt server errors, HTTPS certificate errors, important pages returning errors, and main content missing from the rendered HTML. - Leaks: fix this month. Duplicate versions that don’t redirect, canonicals Google overrides, redirect chains, orphaned service pages, soft 404s and content missing on mobile.
- Hygiene: fix as part of regular care. Sitemap clutter, links to removed pages, parameter duplicates, Core Web Vitals in “needs improvement” and structured data warnings (schema markup for business websites lists the types worth adding).
Plugin updates, new templates and content imports can bring any of these back, so check the Page indexing report monthly and crawl every few months.
What to do next
Start with the blockers: open your robots.txt, run your home page and main service page through URL Inspection, and read every reason in the Page indexing report.
Planning a new website or a redesign? Most of this checklist should pass on launch day, and it’s far cheaper to build in than to retrofit: search blocks removed, a clean sitemap, self-referencing canonicals, full content on mobile and in the HTML, and old URLs redirected in one hop. Does web design include SEO? sets out where a build’s job ends; the website launch checklist orders the go-live checks. Our website design and development projects include clean page structure, an XML sitemap and Search Console set-up, and every website redesign redirects each old URL.
Want a second pair of eyes first? A free website audit needs only your web address, no logins. It checks what Google and a first-time visitor see from outside, including indexing, titles and HTTPS, and shows where on this checklist to start.