Large reference site has real users, but Google excludes 8K+ URLs. What would you prioritize?

barrisb845

Newbie
Joined
Jul 14, 2026
Messages
14
Reaction score
2
I run a network of game-reference wikis across subdomains. GSC reports 8,237 excluded URLs, many as “Crawled - currently not indexed,” even though live tests pass, pages return 200, indexing is allowed, canonicals are self-referencing, and URLs are in sitemaps.

The site has real demand despite limited Google visibility. GA4 for the last seven days shows:

  • 3.8K active users
  • 26K pageviews
  • 2:03 average engagement
  • Major first-user sources: direct (1.5K), Reddit (754), Bing organic (497), and Google organic (313)
The largest individual wiki has about 2,500 structured reference pages and 200 daily unique users.

For a legitimate large reference site, what would you prioritize investigating?
 
I would avoid treating all 8,237 URLs as one indexing problem.

Export them from GSC and group them by subdomain, page template, publication date and internal-link depth. Then inspect a representative sample from each group.

If one template or wiki accounts for most exclusions, compare it for near-duplicate content, orphan depth, crawl frequency, server response time and rendered output.
 
I run a network of game-reference wikis across subdomains. GSC reports 8,237 excluded URLs, many as “Crawled - currently not indexed,” even though live tests pass, pages return 200, indexing is allowed, canonicals are self-referencing, and URLs are in sitemaps.

The site has real demand despite limited Google visibility. GA4 for the last seven days shows:

  • 3.8K active users
  • 26K pageviews
  • 2:03 average engagement
  • Major first-user sources: direct (1.5K), Reddit (754), Bing organic (497), and Google organic (313)
The largest individual wiki has about 2,500 structured reference pages and 200 daily unique users.

For a legitimate large reference site, what would you prioritize investigating?
However, I would first review crawl budget, URL structures, duplicate/ thin content, and internal linking depth. Exclusion of 8k URLs doesn’t mean that there is some technical problem Google just hasn’t found enough value signals on those URLs yet.
 
8k is crazy I’d probably check the crawled but not indexed pages first and see if there’s a pattern. No point trying to fix all 8k individually if the same issue is causing most of them.
 
Back
Top