Skip to main content
mySites.guru
New features added last monthRelease RadarFile ManagerImpostor FilesUpdate QueueRogue AdminsMCP & APIJoomla VELCVE Index

Search engine visibility

What your robots.txt allows search engines and AI answer engines to fetch, and the single stray line that can take a site out of Google altogether.

17 checks in this group.

Joomla

  • Your robots.txt Should Not Block Every Crawler

    Checks whether your robots.txt tells every search engine and AI assistant to stay away from the whole site. This is the single most damaging line a robots.txt can contain, and it is usually left behind by accident after a site goes live.

  • Your robots.txt Should Not Block Search Engines

    Checks whether your robots.txt shuts a conventional search crawler out of the whole site. Blocking Googlebot, Bingbot or Applebot removes you from that search engine, and it is usually collateral damage from a blocklist copied to keep AI scrapers out.

  • Your robots.txt Should Not Block AI Answer Engines

    Checks whether your robots.txt shuts out the crawlers that decide if your site can be cited in an AI answer: OAI-SearchBot, Claude-SearchBot, PerplexityBot and DuckAssistBot. Blocking a training crawler costs you nothing. Blocking one of these deletes you from that assistant.

  • Your robots.txt Should Be Reachable

    Checks that a crawler asking your site for /robots.txt gets a robots.txt back. A 404, a 403, a timeout or an HTML error page all mean search engines are guessing at your rules, and it means we cannot check any of the other findings on this tab.

  • Your robots.txt Must Sit At The Domain Root

    Checks that your robots.txt is where crawlers actually read it. Google reads the file only from the root of each protocol and host, so a site installed in a subfolder has a robots.txt nobody fetches, and the apex and www addresses need a file each.

  • Your robots.txt Should Point Crawlers At Your Sitemap

    Checks whether your robots.txt names your XML sitemap. A declared sitemap is how a crawler finds pages nothing links to, and it is the one line in the file that adds reach rather than removing it.

  • Your AI Training And Answer Engine Policy

    Shows which AI crawlers your robots.txt currently turns away, and separates the ones that cost you nothing from the ones that remove you from ChatGPT, Claude and Perplexity answers.

  • Your robots.txt May Be Managed At The Edge

    Compares the robots.txt on your server with the one crawlers actually receive. When a CDN generates its own, editing the file on disk changes nothing at all.

  • Robots.txt Should Not Block Media & Template Folders

    This is an SEO setting: blocking /media/ or /templates/ in robots.txt stops Google fetching the CSS, JS and images it needs to render your pages properly.

WordPress

  • Your robots.txt Should Not Block Every Crawler

    Checks whether your robots.txt tells every search engine and AI assistant to stay away from the whole site. This is the single most damaging line a robots.txt can contain, and on WordPress it is never something the CMS wrote for you.

  • Your robots.txt Should Not Block Search Engines

    Checks whether your robots.txt shuts a conventional search crawler out of the whole site. Blocking Googlebot, Bingbot or Applebot removes you from that search engine, and it is usually collateral damage from a blocklist copied to keep AI scrapers out.

  • Your robots.txt Should Not Block AI Answer Engines

    Checks whether your robots.txt shuts out the crawlers that decide if your site can be cited in an AI answer: OAI-SearchBot, Claude-SearchBot, PerplexityBot and DuckAssistBot. Blocking a training crawler costs you nothing. Blocking one of these deletes you from that assistant.

  • Your robots.txt Should Be Reachable

    Checks that a crawler asking your site for /robots.txt gets a robots.txt back. A 404, a 403, a timeout or an HTML error page all mean search engines are guessing at your rules, and it means we cannot check any of the other findings on this tab.

  • Your robots.txt Must Sit At The Domain Root

    Checks that your robots.txt is where crawlers actually read it. Google reads the file only from the root of each protocol and host, so a site installed in a subfolder has WordPress generating a robots.txt nobody fetches, and the apex and www addresses need a file each.

  • Your robots.txt Should Point Crawlers At Your Sitemap

    Checks whether the robots.txt crawlers receive names your XML sitemap. A declared sitemap is how a crawler finds pages nothing links to, and it is the one line in the file that adds reach rather than removing it.

  • Your AI Training And Answer Engine Policy

    Shows which AI crawlers your robots.txt currently turns away, and separates the ones that cost you nothing from the ones that remove you from ChatGPT, Claude and Perplexity answers.

  • Your robots.txt May Be Managed Somewhere Else

    Compares the robots.txt on your server with the one crawlers actually receive. On WordPress the file is usually generated rather than stored, and when a CDN generates it instead, editing anything on disk changes nothing at all.

Find out which of these your sites fail

Connect a site and every check in this group runs against it automatically, with the result and the fix in one place. These run twice a day on every connected site.

Run a free audit