WordPress SEO Is a Structural Problem: Taxonomies, Archives, and URL Bloat
WordPress generates pages, archives, and taxonomies by default. Kacper Krasnodębski explains why that default behavior - not bad content - is the root cause of most WordPress SEO problems.

The Problem Is Not the Content
Kacper Krasnodębski opens with a framing that inverts the usual SEO conversation: the content on most WordPress sites is not the problem. The structure around that content is. WordPress, by default, generates a large number of indexable URLs - tag archives, category archives, author pages, date-based archives, paginated listings, filtered product states - and most of them are created without anyone making a deliberate decision. They exist because the CMS makes them easy to produce.
Kacper leads the SEO team at Vilaro, a specialist agency focused on exactly these kinds of structural SEO challenges. If your WordPress site has been publishing content for years but organic performance has plateaued or traffic is fragmented across low-value URLs, Vilaro’s approach - rooted in architecture before tactics - is worth looking at directly. Check their work and see whether your current strategy is working on the right layer.
Tags Are Not a Keyword System
The first structural problem Kacper addresses is the misuse of tags. Because the field in the WordPress editor is labeled “tags” and the concept sounds adjacent to “keywords,” editors use it to store keyword variants: SEO, technical SEO, WordPress SEO, Google ranking, on-page SEO. From an editorial perspective this feels organized. From a structural perspective it creates dozens of indexable URLs with no content depth behind them.
Tags in WordPress are not hidden metadata. They are public taxonomies that generate archive pages. Each new tag - especially if used only once or twice, or named inconsistently across authors - creates a URL that search engines can and will crawl. At scale, a blog with 200 posts and five to ten tags per post can generate hundreds of thin archive pages with no meaningful content cluster behind them.
Kacper’s test for any taxonomy is whether it helps a real user accomplish something. A tag archive is only useful if it aggregates enough related content to justify its existence as a navigational hub. If a tag exists because an editor wanted to signal relevance to a search engine, it should either be consolidated into a meaningful cluster or removed from the index entirely.
The practical prescription: keep the tag system small and intentional. Merge duplicates, remove low-count tags, and consider disabling tag archive indexation if the archives cannot be justified as real content destinations.
Category Archives: High Potential, Low Maintenance
Categories are structurally more important than tags in most WordPress sites. They appear in navigation, breadcrumbs, internal links, and often in URLs. A well-managed category archive can function as a hub page - a strong SEO asset with an introduction, featured articles, and a coherent content cluster behind it.
The problem Kacper describes is that most category archives are created at the start of a project and never revisited. Categories accumulate posts but rarely accumulate editorial attention. Nobody writes introductory copy for the category page itself. The title is whatever was typed at setup. The listing grows but remains a listing.
When a category page lacks unique content, it becomes a thin archive - a paginated list with a bare heading. That page can appear in search results instead of the individual articles that would actually serve the user. It can outrank stronger pages simply by accumulating internal link equity, while delivering little value to anyone who lands on it.
The decision about category archives should be deliberate: if a category represents a real content pillar, invest in it properly - write a useful description, curate what appears there, treat it as a landing page. If a category has no clear user purpose, prevent it from entering the index rather than allowing it to dilute the site’s overall signal quality.
Pagination and the Indexation Question
WordPress generates paginated versions of every archive by default. A category with sixty posts creates page one, page two, page three, and so on. All of these pages are crawlable and indexable unless something explicitly prevents that.
Kacper’s point about pagination is not that it should be blocked universally - deep pagination may still be necessary for discovery. The point is that a user starting their journey on page seven of a category archive is not the intended entry point, and treating those paginated URLs as indexation candidates wastes crawl budget and dilutes the site’s structural clarity.
For content-heavy blogs and WooCommerce stores, this compounds quickly. A product category with a hundred items, three sort orders, and two filter combinations creates a large matrix of crawlable states. Most of those states are useful for navigation but serve no function as search entry points.
The guiding principle Kacper proposes: page one of an important archive can be an SEO target; deeper pagination generally should not be. XML sitemaps should reflect this distinction - they should list URLs you want ranked, not every URL the CMS has generated.
Author Archives and Date Archives: Usually Noise
Kacper identifies two archive types that exist on most WordPress sites without serving any real user purpose: author archives and date-based archives.
Author archives make sense on multi-author publishing platforms where a named contributor has a meaningful content library and the author’s identity is part of the editorial brand. On a single-author company blog or a site where content is produced by a team without public attribution, the author archive is a copy of the main blog listing under a different URL - structural duplication without navigational value.
Date archives are harder to justify on modern sites. Monthly and yearly archive pages were a feature of early blogging culture. On a corporate site or a marketing blog, no user is searching for the May 2023 archive. Yet WordPress can generate and link to those pages automatically, creating more indexable URLs that add noise without adding value.
Both archive types represent the same category of problem: features the CMS enables by default that should require an active decision to keep, not an active decision to disable.
WooCommerce: Where URL Explosion Becomes a Structural Risk
Everything discussed above scales significantly in a WooCommerce context. Product categories, product tags, attribute taxonomies (size, color, material), filter states, sort orders, and pagination all combine to produce a large number of crawlable URL combinations.
Kacper describes this as silent growth. The store looks correct from the front end - products are browsable, filters work, nothing appears broken. But from an SEO perspective, the site is generating thousands of URL variations that no one planned for. Search engines spend crawl activity on those combinations instead of concentrating on category pages and product pages that have real ranking potential.
The structural risk is not just crawl budget in the technical sense. It is also signal dilution: when hundreds of low-value URLs accumulate internal links through automated navigation, the value that should concentrate on important product pages gets dispersed across structural artifacts.
For WooCommerce sites, Kacper’s position is that taxonomy governance cannot be deferred. Filter behavior, canonical handling, and pagination strategy need deliberate configuration - not as an optimization step after launch, but as part of the content model design.
Sitemaps as Structural Signals
Kacper closes with a point about XML sitemaps that follows directly from the structural argument: a sitemap is not passive CMS output. It is a signal. When a sitemap lists tag archives, thin paginated pages, attachment pages, and low-count taxonomies, the site is explicitly telling search engines that these URLs are worth prioritizing.
Most teams treat sitemap generation as automatic - install a plugin, configure the basics, forget it. The result is a sitemap that reflects the CMS’s default output rather than the site’s actual SEO priorities. A selective sitemap - one that includes only URLs with genuine indexation value - reinforces the structural clarity that the rest of the configuration is trying to achieve.
Sitemap management should follow structural decisions, not replace them. It is the last step in a coherent indexation strategy, not a shortcut around it.
Structural Discipline Over URL Proliferation
Kacper’s closing argument is precise: WordPress is not inherently bad for SEO, but unmanaged WordPress is very good at creating SEO problems. The CMS generates pages, archives, and taxonomies easily. That ease is also the risk. Without an active structural audit - asking which URLs serve a real user purpose and which exist only because the CMS defaults produced them - a WordPress site accumulates a growing index of low-value pages that dilute the signal quality of the content that actually matters.
The question to ask is not only “what content do we have?” but “what has the system built around that content, and does any of it belong in search?”
This is exactly the kind of structural SEO work that Vilaro specializes in. If your WordPress or WooCommerce site is generating more URLs than your team has ever audited, the Vilaro team can help you identify what is working against your rankings and what a governed structure looks like in practice. The performance gap in structural SEO is not usually a content problem - and fixing it starts with knowing where the problem actually lives.
Tags
Keep Reading

Or maybe a warm-up before Gdynia? Join us at CMS Conf Local in Vilnius
Before we meet at the main CMS Conf 2026 in Gdynia on November 12–14, we’ll see you on October 24 at CMS Conf Local: Vilnius.

From Browsers to Buyers
Katarzyna Janoska on why generic copy kills conversion faster than slow load times, and why reassurance is the real job of a small business website.

Most AI Tools Won't Survive
Damian Ślimak on twenty years of automation before it had a name, why clients never ask whether AI is involved, and the one skill still worth training.