> Chris Abraham manages crawling and indexing on large and programmatic sites: segmented sitemaps, parameter URLs, index bloat, internal linking, and quality.
>
> Source: https://gerriscorp.com/services/large-sites/ · Updated 2026-10-05 · By Chris Abraham, Gerris Corp

# SEO for large and programmatic sites

Sites with tens of thousands of pages, built from databases, feeds, or templates, live or die on how search engines crawl and index them. The work is managing collections of pages, not tuning a handful.

## The usual story

A programmatic site launches, indexing climbs fast, and then a large share of pages drops out. Search Console fills with "Crawled, currently not indexed" and "Discovered, currently not indexed." The cause is rarely one bug. It's usually some mix of near-duplicate pages, URL variations generated by tools and filters, weak internal linking to deep pages, thin templates, and, on expired or repurposed domains, history the site inherited.

## What I work on
- **Segmented XML sitemaps.** Separate sitemaps for each page type or collection, so Search Console reports indexing coverage per segment and problems show up where they live.
- **Baselines and comparisons.** Indexing measured per segment over time, so changes are judged against data rather than impressions.
- **Parameter and tool URLs.** Calculators, filters, sort orders, and search tools that spawn endless URL variations, handled with robots rules, canonicals, and link changes.
- **Index bloat.** Deciding which page types deserve indexing at all, and which should be merged, improved, or kept out.
- **Template quality.** Making sure each generated page carries enough unique, useful information to earn its place, rather than a template with a name swapped in.
- **Internal linking at scale.** Hub pages, related-item links, and breadcrumbs that give deep pages a path from the home page.
- **Crawl efficiency.** Server response times, crawl stats, redirect chains, and soft 404s that waste crawl attention.
- **Reporting lag.** Search Console reports trail reality by days or weeks. I check live behavior and server evidence before concluding a fix failed or worked.

## Working with the development team

Large sites are built by developers, so the fixes go through them. I work in the repository and the team's chat tools, write changes as specs with acceptance criteria, and verify them with crawls and Search Console data. See [developer specs and QA](https://gerriscorp.com/services/developers/).

## Common questions

### Why did my programmatic pages drop out of the index?

Google indexes generously at first and then reassesses. Pages that look alike, add little beyond a template, or sit far from internal links are the first to drop. A diagnostic finds which segments are affected and why.

### Does an expired domain help or hurt?

It can do either. Old links can help a new site get noticed, but a domain's history and a sudden change of topic can also weigh on how Google treats it. It's one more factor to separate from the site's own problems.

**Related:** ["Crawled, currently not indexed"](https://gerriscorp.com/guides/crawled-not-indexed/) · [Search Console](https://gerriscorp.com/services/search-console/) · [E-commerce SEO](https://gerriscorp.com/services/ecommerce/)

[Tell me what dropped out](https://gerriscorp.com/contact/)

Updated October 5, 2026
