[wtf] Mysterious Homepage De-Indexing

ipancake

Our homepage, as well as several similar landing pages, have vanished from the index. Could you guys review the below pages to make sure I'm not missing something really obvious?!

URLs: http://www.grammarly.com http://www.grammarly.com/plagiarism-checker

It's been four days, so it's not just a temporary fluctuation
The pages don't have a "noindex" tag on them and aren't being excluded in our robots.txt
There's no notification about a penalty in WMT

Clues:

WMT is returning an "HTTP 200 OK" for Fetch, is showing a redirect to grammarly.com/1 (alternate version of homepage, contains rel=canonical back to homepage) for Fetch+Render. Could this be causing a circular redirect?
Some pages on our domain are ranking fine, e.g. https://www.google.com/search?q=grammarly+answers
A month ago, we redesigned the pages in question. The new versions are pretty script-heavy, as you can see.
We don't have a sitemap set up yet.

Any ideas? Thanks in advance, friends!

Dr-Pete

Did this get resolved? I'm seeing your home-page indexed and ranking now.

I'm not seeing any kind of redirect to an alternate URL at this point (either as a browser or as GoogleBot). If you 301'ed to an alternate URL and then rel=canonical'ed back to the source of the 301, that could definitely cause problems. It's sending a pretty strong mixed-signal. In that case you'd probably want to 302 or use some alternate method. Redirects for the home-page are best avoided, in most cases.

RyanPurkey

Are you sure it was missing for a time? Ultimately I wouldn't use a third-party (Google) as a tool to diagnose problems (faulty on-site code) that I know are problems and need to be fixed.I'd fix the problems I know are issues and then go from there. Or hire someone capable of fixing the problems.

ipancake

Thanks, Ryan. I'll get to work on the issues you mentioned.

I do have one question for you - grammarly.com/proofreading (significantly fewer links, identical codebase) is now back on the index. If the issue was too many scripts or HTML errors, wouldn't both pages still be de-indexed?

RyanPurkey

Here are some issues just going down the first few lines of code...

There's a height attribute in your tag.
Your cookie on the home page is set to expire in the past, not the future
Your tag conflicts with your script and other code issues (http://stackoverflow.com/questions/21363090/doctype-html-ruins-my-script)
Your Google Site Verification meta tag is different than other pages.
Your link to the Optimizely CDN is incorrect... (missing 'http:' so it's looking for the script on your site)
You have many other Markup Issues.

And that's prior to getting into the hundreds of lines of code preceding the start of your page at the tag... 300 lines or so on your other indexed pages 1100+ on your home page. So not only are you not following best practices as outlined by Google, but you have broken stuff too.

ipancake

The saga continues...

According to WMT, there are no issues with grammarly.com The page is fetched and rendered correctly.

Google! Y u no index? Any ideas?

RyanPurkey

Like Lynn mentioned below, if you're having redirection take place across several portions of the site, that could cause the spikes, and a big increase in total download time is worrying if you're crossing the average bounce rate threshold for most people's patience.

Here's the Google Page speed take on it: https://developers.google.com/speed/pagespeed/insights/?url=http%3A%2F%2Fgrammarly.com&tab=desktop. They go over both desktop and mobile.

LynnPatchett

Hmm, was something done to fix the googlebot redirect issue or did it just fix itself? Here it states that googlebot will often identify itself as mozilla and your fetch/render originally seemed to indicate that at least some of the time that was the page google was getting. It is a bit murky technically what exactly is going on there but if google is getting redirected some of the time then as you said you are getting into a circular situation between the redirect and the canonical where it is a bit difficult to predict what will happen. If that is 100% fixed now and google sees the main page all the time then I would wait a day or two to see if the page comes back into the index (but be 100% sure that you know it is fixed!). I still think that is the most likely source of your troubles...

ipancake

Excellent question, Lynn. Thank you for chiming in here. There's a user agent based javascript redirect that keeps Chrome visitors on grammarly.com (Chrome browser extension) and sends other browsers to grammarly.com/1 (Web app that works on all browsers).

UPDATE: According to WMT Fetch+Render, the Googlebot redirection issue has been fixed. It is no longer being redirected anywhere and returning a 200 OK for grammarly.com.

Kelly, if that was causing the problem, how long should I hold my breath for re-indexing after re-submitting the homepage?

Anti-Alex

Yup definitely. Whether you're completely removed or simply dropped doesn't matter. If you're not there anymore, for some reason Google determined you're no longer an authority for that keyword. So you need to find out why. Since you just redesigned, the way way is to back track, double check all the old tags and compare them to the new site, check the text and keyword usage on the website, look for anything that's changed that could contribute to the drop. If you don't find anything, tools like majesticSEO are handy to checking if your backlinks are still healthy.

ipancake

Hi Alex, Thank you for your response. The pages didn't suffer in ranking, they were completely removed from the index. Based on that, do you still think it could be a keyword issue?

ipancake

That's actually a great point. I suppose Google could have been holding on to a pre-redesign cached version of the pages.

There has been a 50-100% increase in page download times as well as some weird 5x spikes for crawled pages. I know there could probably be a million different reasons, but do any of them stick out at you as being potential sources of the problem?

LynnPatchett

How does that second version of the homepage work and how long has it been around for? I get one version of the homepage in one browser and the second in another, what decides which version is served and what kind of redirect is it? I think that is the most likely source of your troubles.

RyanPurkey

Yes, but the pages were indexed prior to the redesign, no? Can you look up your crawl stats in GWT to see if there's been a dramatic up tick in page download times, and a down trend in pages crawled. That will at least give you a starting point as to differences between now and then: https://www.google.com/webmasters/tools/crawl-stats

Anti-Alex

Logo definitely needs to be made clickable to Home.

Did you compare the old design and the new design's text to make sure you're still covering the same keywords. In many cases a redesign is more "streamlined" which also means less text or a re-write which is going to impact the keywords your site is relevant for.

ipancake

Thanks, Ryan. Improving our code-to-text ratio is on our roadmap, but could that really be the issue here? The pages were all fully indexed without problems for a full month after our redesign, and we haven't added any scripts. Was there an algorithm update on Monday that could explain the sudden de-indexing?

RyanPurkey

VERY script heavy. Google has recently released updates on a lot of this (Q4 2014) here: http://googlewebmastercentral.blogspot.mx/2014/10/updating-our-technical-webmaster.html. With further guidance given here: https://developers.google.com/web/fundamentals/performance/optimizing-content-efficiency/optimize-encoding-and-transfer. Without doing a deep dive that's the most glaring issue and obvious difference between pages that are still being indexed and those that are not.

Welcome to the Q&A Forum

Browse the forum for helpful insights and fresh discussions about all things SEO.

[wtf] Mysterious Homepage De-Indexing

Browse Questions

Explore more categories

Related Questions

Google Indexing Request - Typical Time to Complete?

Case Study Mystery

Using del canonical for subpage relating to homepage

Old pages STILL indexed...

Content From One Domain Mysteriously Indexing Under a Different Domain's URL

How is Google crawling and indexing this directory listing?

How to deal with old, indexed hashbang URLs?

Should I Allow Blog Tag Pages to be Indexed?