How To Set Up a Proper Robots.txt for WordPress | Control Search Indexing, Robots.txt Guide | SEOShareLink
Ever wondered if you’re accidentally putting up invisible “Keep Out” signs for Google in the back alleys of your WordPress site?
You are. It’s called your robots.txt file. This small, powerful text file is your site’s primary set of instructions for search engine crawlers. In WordPress, a misconfigured one can block Google from seeing your most important pages. Setting it up correctly is one of the smartest 10-minute SEO tasks you can do.
Let’s build yours the right way.
The Robots.txt File: Your Crawler Traffic Director
Think of your website as a museum. The robots.txt file is the guide at the entrance, politely telling the search engine bots which wings are open for public tours and which storage rooms are for staff only.
- Its Job: To guide. It uses simple commands to allow or disallow crawlers from accessing certain folders and files.
- A Key Limitation: It’s a request, not a law. Malicious bots can ignore it. For absolute blocking, you need stronger methods like password protection or a
noindexmeta tag.
“A well-crafted
robots.txtfile isn’t about being secretive; it’s about being an efficient host. You’re helping Google’s crawler spend its valuable crawl budget on your public content, not wasting time in private areas.”
The WordPress-Specific “Gotchas” You Must Know
WordPress automatically creates a basic robots.txt file, but it’s often incomplete or hidden. Knowing its quirks is half the battle.
- The “Virtual” File: WordPress often generates a virtual
robots.txtfile on the fly. To see it, go tohttps://yoursite.com/robots.txt. You may need to adjust your permalink settings for it to appear. - The Hidden Admin Area: The most important thing you should block is your
/wp-admin/directory. You don’t want Google trying to index login pages or backend scripts. - Plugin Clutter: Many SEO plugins (like Yoast) will manage this file for you, but they can sometimes add overly complex or conflicting rules.
Building Your Optimized Robots.txt File (Step-by-Step)
Let’s create a clean, effective file. You can directly edit it if it exists, or create it.
Step 1: Access & Edit the File
The most reliable method is to create or edit the physical file via FTP or your hosting file manager.
- Connect to your site via FTP (using FileZilla or similar).
- Navigate to the root directory of your website (often
public_html,www, or named after your domain). - Look for a file named
robots.txt. If it exists, download a backup copy to your computer first. If it doesn’t, create a new empty text file and name it exactlyrobots.txt. - Open the file in a plain text editor (like Notepad or TextEdit).
Step 2: Paste the Optimized WordPress Template
Copy and paste the following code. This is a standard, safe, and SEO-friendly starting point for most WordPress sites.
User-agent: *
Allow: /wp-admin/admin-ajax.php
Allow: /wp-content/uploads/
Disallow: /wp-admin/
Disallow: /wp-includes/
Disallow: /wp-login.php
Disallow: /wp-signup.php
Disallow: /readme.html
Disallow: /license.txt
Disallow: /xmlrpc.php
Disallow: /*?*
Disallow: /search/
Sitemap: https://www.yoursite.com/sitemap_index.xml
Step 3: Understand the Rules (What You’re Blocking and Why)
User-agent: *: This rule applies to all respectful crawlers (Googlebot, Bingbot, etc.).Allow: /wp-admin/admin-ajax.php: A crucial exception! This allows scripts needed for your site’s functionality (like forms) to work. Do not omit this.Allow: /wp-content/uploads/: This explicitly allows crawlers to access your images and media files so they can be indexed.Disallow: /wp-admin/,/wp-includes/, etc.: These block crawlers from your WordPress core files, login pages, and other non-public areas. This is good hygiene.Disallow: /*?*: Blocks URLs with query strings (like?s=searchor?p=123), which often create duplicate content.Disallow: /search/: Stops Google from indexing internal search result pages, which are low-value duplicates.Sitemap:This line is vitally important. It tells crawlers exactly where to find your sitemap. You must replace the example URL with your actual sitemap URL (often/sitemap.xmlor/sitemap_index.xml).
Testing & Validation: Don’t Just Set and Forget
After uploading your new robots.txt file back to your site’s root, you must test it.
- View It: Go to
https://yoursite.com/robots.txtin your browser. You should see your new file. - Test with Google: Use Google Search Console’s Robots.txt Tester.
- Go to Settings > Robots.txt Tester in the left menu.
- Click “Test” to see if Google can fetch it and if there are syntax errors.
- Use the “Allow” and “Block” indicators at the bottom to test specific URLs on your site to see if they would be blocked.
- Resubmit Your Sitemap: In Google Search Console, go to Sitemaps and resubmit your sitemap to ensure Google picks up the new path from your
robots.txt.
Your WordPress Robots.txt FAQ
What’s the difference between robots.txt and a noindex meta tag?robots.txt says, “You are not allowed to come into this room.” It controls access for crawling. A noindex meta tag says, “You can come in, but don’t put this in your index/show it in search results.” For blocking sensitive pages (like admin), you should use both: block crawling and prevent indexing.
Should I block /wp-content/themes/ or /wp-content/plugins/?
It’s generally safe to block these (Disallow: /wp-content/themes/, Disallow: /wp-content/plugins/). They often contain files that aren’t meant for public viewing. However, some themes/plugins serve public CSS or JS from there. If you add these rules and your site looks broken, remove them.
My site uses query strings for filters (e.g., ?color=blue). Won’t Disallow: /*?* block those?
Yes, it will. If you have important filtered pages you do want indexed (common in e-commerce), this rule is too broad. You’ll need a more sophisticated robots.txt that uses specific Allow rules to override the Disallow. This is a case where a plugin or developer help might be needed.
What if my WordPress SEO plugin is creating its own robots.txt?
If you’re using a plugin like Yoast, Rank Math, or All in One SEO, it likely has a dedicated settings panel for robots.txt. You should use that panel instead of a physical file, as the plugin will manage it. Do not have both a physical file and plugin-generated rules—they will conflict.
How do I block a specific search engine bot, like a scraper?
You can target specific user-agents. For example, to block a common content scraper, you could add:
User-agent: AhrefsBot
Disallow: /
You can find lists of common bot user-agent names online.
Ready to give Google the perfect map of your site? Your first action is the easiest: Go to yoursite.com/robots.txt right now. What do you see? A simple, clean file like the one above? Or a mess of conflicting rules? Just knowing its current state is the first step to taking control.