Will disallowing URL's in the robots.txt file stop those URL's being indexed by Google

andyheath

I found a lot of duplicate title tags showing in Google Webmaster Tools. When I visited the URL's that these duplicates belonged to, I found that they were just images from a gallery that we didn't particularly want Google to index. There is no benefit to the end user in these image pages being indexed in Google.

Our developer has told us that these urls are created by a module and are not "real" pages in the CMS.

They would like to add the following to our robots.txt file

Disallow: /catalog/product/gallery/

QUESTION: If the these pages are already indexed by Google, will this adjustment to the robots.txt file help to remove the pages from the index?

We don't want these pages to be found.

Martijn_Scheijbeler

That's why I mentioned: "eventually". But thanks for the added information. Hopefully it's clear now for the original poster.

andyheath

Looking at this video - https://www.youtube.com/watch?v=KBdEwpRQRD0&feature=youtu.be Matt Cutts advises to use the noindex tag on every individual page. However, this is very time consuming if you're dealing wit a large volume of pages.

The other option he recommends is to use the robots.txt file as well as the URL removal tool in GWMT, Although this is the second choice option, it does seem easier for us to implement than the noindex tag.

varun1800

Hi,

Yes, if you put any url in the robots.txt it will not be shown in the search results after some time even if your pages were already indexed. Because when your disallow urls in the robots.txt , Google will stop crawling that page and eventually will stop indexing those pages.

andyheath

Hi Nico

Great response thanks.

This is certainly something I'm taking into consideration and will question my developer about this.

andyheath

Thanks Thomas.

I'm now finding out from my developer is we are able to noindex these pages with the meta robots.

If this is something that isn't possible, it's likely that we'll add to the robots.txt as you did.

Either way I think will be progress to different degrees.

ThomasHarvey

I don' think Martijn's statement is quite correct as I have made different experiences in an accidental experiment. Crawling is not the same as indexing. Google will put pages it cannot crawl into the index ... and they will stay there unless removed somehow. They will probably only show up for specific searches, though

Completely agree, I have done the same for a website I am doing work with, ideally we would noindex with meta robots however that isn't possible. So instead we added to the robots.txt, the number of indexed pages have dropped, yet when you search exactly it just says the description can't be reached.

So I was happy with the results as they're now not ranking for the terms they were.

netzkern_AG

I don' think Martijn's statement is quite correct as I have made different experiences in an accidental experiment. Crawling is not the same as indexing. Google will put pages it cannot crawl into the index ... and they will stay there unless removed somehow. They will probably only show up for specific searches, though

In September 2015 I catapulted a website from ~3.000 to 130.000 indexed pages (roughly). 127.000 were essentially canonicalised duplicates (yes, it did make sense) but also blocked by robots.txt - but put into the index nonetheless. The problem was a dynamically generated parameter, always different, always blocked by robots.

The title was equal to the link text; the description became "A description for this result is not available because of this site's robots.txt – learn more." (If Google cannot crawl a URL Google will usually take titles from links pointing to that URL). No sign of disappearing. In fact, Google was happy to add more and more to its index ...

At the start of December 2015 I removed the robots.txt block - Google could now read the canonicals or noindex on the URLs ... the pages only began dropping out, slowly and in bunches of a few thousand in March 2016 - probably due to the very low relevancy and crawl budget assigned to them. Right now there are still about 24.000 pages in the index.

So my answer would be: No - disabling crawling in the robots.txt will NOT remove a page from the index. For that you need to noindex them (which sometimes also works if done in robots.txt, I've heard). Disallowing URLs in the robots.txt will very likely drop pages to the end of useful results, though, as Andy described. (I don't know if this has any influence on the general evaluation of the site as a whole; I'd guess not.)

Regards

Nico

andyheath

Thanks Martijn. This is what I was assuming would happen. However, I got a confusing message from my developer which said the following,

"won't remove the URL's from the index but it will mean that they will only show up for very specific searches that customers are extremely unlikely to use. It will also increase Asgard's crawl budget as Google and Bing won't try to crawl these URLs. Would you be happy with this solution?"

I would tend to still agree with your statement though.

Martijn_Scheijbeler

Yes they will be eventually. As you disallow Google to crawl the URLs it will probably start hiding the descriptions for some of these image pages soon as they can't crawl them anymore. Then at some point they'll stop looking at them at all.

Welcome to the Q&A Forum

Browse the forum for helpful insights and fresh discussions about all things SEO.

Will disallowing URL's in the robots.txt file stop those URL's being indexed by Google

Browse Questions

Explore more categories

Related Questions

Best practice for disallowing URLS with Robots.txt

"Null" appearing as top keyword in "Content Keywords" under Google index in Google Search Console

How to make Google index your site? (Blocked with robots.txt for a long time)

Https & http urls in Google Index

Pages getting into Google Index, blocked by Robots.txt??

Best way to permanently remove URLs from the Google index?

Wordpress blog in a subdirectory not being indexed by Google

Indexed non existent pages, problem appeared after we 301d the url/index to the url.