Crawling Question -- regarding "Excluded" in the Coverage Reports in GSC

KurtTeej

Newbie
Joined
Jul 22, 2019
Messages
15
Reaction score
2
I've got a client that recently did what I'll call a "refresh". When this client described the project, they described it as a redesign, but they also rolled out a new method of building the pages rather than just delivering straight html code.

I've been monitoring the Coverage reports to make sure that the site has been fully crawled as it's a fairly large site (175,000 or so pages). One of the 4 reports included in GSCs Coverage reports is the "Excluded" report. This report essentially lists urls that the crawler encounters that has a proper canonical tag so Google has excluded the url from the index. In most cases, I'd do a cursory review and be roughly okay with most of what's in there because the urls ARE in fact okay and the canonical points to the correct page. For example:
https://sitename.com/index.html?track=menu
This url is one the site to provide some tracking information to illustrate to someone that someone selected the home page from something called "menu" (OVER SIMPLIFIED EXAMPLE), so something like this is perfectly normal and fine that it gets listed as "Excluded".

Here's the problem. Since this project was launched (end of April), the "Excluded" bucket has absolutely exploded --> June 1: 14.5 million, June 13: 26.9 million, June 20: 36.3 million, June 27: 49.6 million. In the list of examples, the urls that are presented (they only give you access to 1,000) are all what I would call "Junk", they are not real urls with some parameter for tracking or anything else.

Point #2 - looking at the old Crawl Stats charts, the number of pages crawled every day grew throughout June but has topped out at 2.4-2.6 million pages. The time spent downloading a page has grown from .2 seconds up to 3.9 seconds in the latest data.

My point to the client is that this is a problem that has to be addressed and fixed because more and more of the time that Google has allotted to this client is being taken up looking at pages that aren't really pages and don't exist.

Am I being obsessive-compulsive on this or am i correct in my assertion that this is in fact a pretty big problem that needs to get fixed sooner rather than later? [additional note --> rankings have dropped pretty much across the board since this 'refresh' was launched and organic traffic has been cut significantly.]
 
Back
Top