Faceted navigation can create thousands of near-duplicate URLs
Why a useful property filtering interface can create an enormous crawl space, and why the solution begins with deciding which filter states deserve to exist in search.
This article reflects the named expert’s practical perspective. See NovAsia’s editorial policy for how material is prepared and reviewed.
Faceted navigation solves a real property-search problem. A buyer wants to combine location, budget, bedrooms, completion status and other preferences without learning the site's taxonomy first. The interface should let them do that. The SEO problem begins only when every possible combination also becomes a crawlable address that looks like a new page.
This can happen quietly. A catalogue with hundreds of projects may expose many thousands of URL combinations once parameters, sorting, pagination and alternate parameter orders are included. The underlying inventory has not multiplied. The address space has.
The product needs freedom; the index needs a policy
Google's current crawling documentation is unusually direct about faceted navigation: parameter-based combinations can generate an effectively unbounded URL space. Crawlers may spend substantial resources visiting combinations that do not provide useful search destinations, which can slow discovery of new, useful URLs.
That warning should not be translated into “filters are bad for SEO.” Filters are often excellent for people. The correct translation is that a site needs a deliberate distinction between interactive states and search destinations.
Take a property catalogue where a user can select Phuket, two bedrooms, a maximum price, sea view and completed status. That combination may be completely reasonable for one browsing session. It does not follow that the site needs a permanent indexed page for the combination. Inventory can change tomorrow, and the same state might contain five projects, one project or none. If the URL exists only to preserve an interface selection, treating it as durable editorial content can create more maintenance than value.
The policy becomes clearer once stable destinations are listed first. City catalogues may deserve permanent URLs. Some district or completion-status pages may deserve them too if they represent distinct, maintained user jobs. Arbitrary combinations can still work perfectly inside the product without joining the same indexable hierarchy.
Technical controls should follow the content model
The technical layer becomes much easier to reason about once those destinations are named. A crawler rule, canonical signal or indexability setting then implements a decision the team can explain, instead of becoming a substitute for that decision. This matters because several individually reasonable signals can create confusion when they point in different directions.
Technical signals come after that decision. Google documents different approaches depending on whether faceted URLs are intended to be crawled and indexed. If they are not useful search destinations, crawling can be constrained. If some need to remain discoverable, their URL behaviour should be consistent and their empty or nonsensical states handled properly. The important part is that the implementation follows a content model instead of trying to invent the model with tags.
Canonicalization deserves the same restraint. Google lists filtering and sorting among common causes of duplicate URLs and uses canonicalization to select a representative version. But a canonical hint does not turn an uncontrolled URL generator into a clear architecture. Internal links, sitemaps, parameter behaviour and indexability should not point in contradictory directions.
For a large property site I would therefore ask a nontechnical question before touching configuration: if this filtered URL appeared as a search result, would a visitor reasonably see it as a stable answer in its own right? If the answer is no, the search system probably does not need every variation even though the catalogue interface does.
That test also protects against a different mistake: automatically blocking every facet. Some combinations may become genuinely useful landing pages because they represent recurring, distinct choices. The difference is not the syntax of the URL. It is whether the page has a maintained purpose, meaningful inventory and an answer that remains useful beyond one session.
The system is strongest when those two layers are allowed to do different jobs. The buyer gets flexible exploration. Search gets a finite set of destinations the site is prepared to maintain and explain. Trying to make every filter state serve both purposes is what creates the near-duplicate maze.
Sources
- Google Crawling Infrastructure — **Managing crawling of faceted navigation URLs**. Confirms that faceted navigation can generate extremely large URL spaces, cause overcrawling, and slow discovery of useful new URLs; it also distinguishes approaches for facets that do or do not need search visibility. Accessed 2026-10-06.
- Google Search Central — **What is canonicalization**. Confirms that sorting and filtering are common sources of duplicate URLs and explains canonical selection as deduplication. Accessed 2026-10-06.