What is Index Bloat? — Whiteboard Friday - Moz

What is Index Bloat? — Whiteboard Friday

May 8, 2026 9 min read

Written by: Tom Capper Edited By: Meghan Pahinui

Table of Contents

Is your website carrying deadweight? Index bloat often affects medium to large websites. Learn how to identify URLs consuming your index quota without delivering traffic, understand the difference between crawl budget and index bloat, and discover practical solutions for cleanup. This Whiteboard Friday video helps you assess your site's index health and implement effective remediation steps, from content consolidation to proper URL handling. Because if you don't trim the excess pages from your site, you're making Google work harder to find—and trust—your best content. Let's fix that.

This popular episode was originally published in February 2025 and still provides a lot of value to SEOs and marketers.

Understanding index bloat

But before we get into all of that, I just want to explain this. So I put this diagram in just to give a bit of context, so that what I say next makes sense. So this outer box, this diagram as a whole is all of the URLs on your site, all of the URLs that might exist, that could exist, including parameters that no one has tried before, this kind of thing, the maximum possible set of URLs that would return a 200 response code and a valid page.

And then I've got smaller sets within that, sort of subsets of URLs. So the next one down is Google discovered URLs. So if Google has seen the URL – they might have not crawled it, they might have not indexed it, but they've seen the URL, they know it exists. That's sort of your next step down. And if you've got a big difference between the red box and the blue box, that probably indicates some kind of crawl budget problem. But that's not what we're talking about today.

If you've got a URL that's discovered, it might not necessarily be indexed. So indexed URLs is another smaller set. If you've got a URL that's discovered but not indexed, again there might be some reasons for that. Google might suspect that the page is unimportant based on other signals.

And then you've got indexed versus pages with non-trivial traffic. Now what counts as non-trivial traffic to a page might vary from site to site. You might have your own idea of this. But a big gap between the number of URLs that are indexed and the number of URLs that are getting any kind of meaningful non-zero traffic, if that's a big gap, that would suggest an index bloat problem, and that's what I want to talk about today.

What index bloat is not

So before I get into that, just to make it totally clear, I want to quickly disambiguate a couple of things I mentioned there. So I'm not talking about crawl budget. As I mentioned before, that's when you've got a lot of URLs that Google just isn't going to crawl at all.

I'm also not talking about cannibalization. Now that is a related concept. Often when you've got a huge number of indexed pages that aren't getting traffic, it's because their topics are too similar. But you could theoretically have a cannibalization problem on a site with three pages, if they're all about roughly the same thing. That's not really what I'm talking about today. I'm talking about a larger-scale problem.

Why is index bloat an issue?

Why do we care? Why is that a problem? So what? I've got lots of indexed pages that don't get any traffic. What's the big issue?

Well, the first thing is that we have to theorize about how Google sort of treats these and why it behaves the way it does, and why we see the results we do. This is mostly based on experience within the industry. It's not something that Google has ever sort of codified for us.

The other is that, as I just mentioned, cannibalization, but also some other technical SEO problems. This can be symptomatic of other issues. If you're generating these large numbers of URLs on your site that are getting indexed, if we think about sort of Page Rank in a very old-fashioned SEO way, that's creating a lot of loss as Google sort of dilutes Page Rank across all of these pages on your site that could be consolidated into the pages that actually have potential to deliver traffic.

What are some common causes of index bloat?

Some common causes we might have. So if we've got all of these URLs that are indexed and not receiving traffic, in a lot of sites that couldn't happen, right?

If you have an editorial policy where you're constantly reviewing and creating pages based on demand, you would hope that this just wouldn't happen.

But on a lot of sites, it does happen, and there are two sort of groups of common reasons for that.

The second sort of group to my mind is listings or products. So if you imagine a real estate website or a used car website or a job listings board, anything like this, a marketplace, you're going to get a lot of these pages that just come and go.

How to reduce index bloat

So in both of these cases, you can generate this very large number of URLs that are getting basically no traffic. So what might we actually do about this or decide whether we want to do something about this?

1. Identify URLs with almost no traffic

So the first step I would take is identify URLs which have near as dammit no traffic. And a rule of thumb I've often used in the past is are they getting on average less than one click a month or something like this?

2. Improve any pages that are opportunities

Next up, improve any that are opportunities. So that's kind of a big catchall statement. But if some of these pages, that you identify, perhaps you used to get a lot of traffic but have become out of date or something like that, or you think they do actually have quality content on them, maybe there's a technical SEO issue that's holding them back.

3. Consolidate or cull pages you’re not able to improve

And then what's left, you've got this big bunch of pages that get zero traffic that you don't think are any kind of opportunity.

So for anything where you really just don't serve this intent, it's redundant, it never had any value anyway, you can just 404 or noindex.

So yeah, this is the sort of process I've followed myself in the past. It's something I've seen good results with. I've seen a lot of other SEOs speaking about this, especially in the wake of the helpful content update and in the past around Panda, which I think probably worked quite similarly.

Let me know how you get on. And thank you very much.

The author's views are entirely their own (excluding the unlikely event of hypnosis) and may not always reflect the views of Moz.