Magento 2 technical SEO audit: URLs, robots, hreflang and JSON-LD
When important Magento 2 products are difficult to find, start with their URLs, content and publishing process. An additional extension cannot correct conflicting directives, outdated claims or a broken code example by itself. This checklist helps identify the cause, choose the next action and verify the result.
1. Choose a sample around the problem
Inspect a simple product, configurable product, category, CMS page and article. Include equivalent pages in another language. Record the date, HTTP status, final URL, title, robots directive and canonical. Separate a technical fault from missing evidence of indexing: a 200 response does not prove that a page is indexed.
2. Compare crawler access with page directives
Check robots.txt alongside the robots metatag and X-Robots-Tag header. Robots.txt is neither an access control system nor a substitute for noindex. If you intend to exclude a page from search, ensure the crawler can read the relevant directive. The Google robots.txt specification explains crawler group matching.
3. Align preferred URLs and languages
Compare canonical, internal links, redirects and XML sitemaps. Inspect parameter URLs, category paths and old domains. Canonical and hreflang have different purposes. Hreflang can be supplied through XML sitemaps without duplicating it in HTML; validate reciprocal and consistent links. See Google’s localized page guidance.
4. Validate product data and JSON-LD
Compare names, prices, currencies, availability, images and reviews with the visible offer. Check whether the theme and multiple extensions emit conflicting data. Use a JSON parser and Rich Results Test: valid syntax alone does not establish content accuracy. Markup does not guarantee rich results. Escape code examples so that a literal <script> in an instruction cannot execute as a real script. Google structured data policies.
5. Make answers and feature boundaries clear
Ground each FAQ answer in that product’s documentation. Distinguish a delivered feature from a separate integration, supported requirements from a tested configuration, and a possible benefit from a guaranteed outcome. Correct useful inaccurate answers and withdraw duplicates. Google stopped displaying FAQ rich results on May 7, 2026; useful answers still help customers. Google Search documentation updates.
6. Verify exports and AI crawler access
Compare published content with its feed: URL, language, status, update time and unique identifiers. Verify export refresh after edits. For ChatGPT Search, check OAI-SearchBot rules and access from published IP ranges using the OpenAI crawler documentation. A request bearing the same user agent from your computer does not prove access from the crawler’s network. Neither llms.txt nor a feed guarantees citations or recommendations.
Match a tool to the confirmed issue
- Kowal Sitemap — XML sitemap coverage and generation.
- SEO Rich Data — structured data that must match the visible offer.
- Kowal Robots — CMS page robots metatags, rather than the site’s robots.txt configuration.
- Canonical URLs — CMS and custom page canonicals, rather than universal product canonical configuration.
Close the audit with evidence
Record the before and after state, change operation ID and readback result. Check the public page after cache refresh. Measure visibility separately with a fixed query set, date, cited sources and the service tested. A successful content save, search indexing and an AI mention are three different outcomes.