# SEO

SEO Site Architecture: Page Maps, Crawl Paths and Governance

SEO Site Architecture: Page Maps, Crawl Paths and Governance

SEO site architecture is the operating model behind the website: which pages exist, what job each URL owns, how people and crawlers reach them, which variants may be indexed and how the structure remains coherent as the business changes. It is broader than a menu and cannot be reduced to a folder pattern or a fixed click-depth score.

The governing rule

Create one durable preferred page for each necessary user task, then connect it through crawlable journeys, consistent index signals and an owner who can keep it accurate.

Define architecture as a set of controlled decisions

LayerDecisionRequired output
AudienceWho must complete which task?Journey and evidence map
Page systemWhich URL owns each task?Approved page map
HierarchyHow are pages grouped?Parent, child and sibling relationships
DiscoveryHow are pages reached?Navigation and contextual link paths
Index controlWhich versions may appear in Search?Canonical, noindex, robots and status rules
OperationsWho approves change?Owner, trigger and QA record

Give every page one primary job

A page may support several secondary actions, but it needs one primary audience, intent and outcome. When two URLs own the same job, they compete for maintenance, links, measurement and user attention even if Google does not treat them as a ranking problem.

Use search intent and first-party customer questions to define the job before writing a slug.

Page jobPrimary outcomeTypical evidence
HomepageOrient and routeClear positioning and major choices
HubChoose a path within a subjectCurated child resources
Service/productEvaluate an offerScope, fit, proof and next step
GuideComplete a learning or operating taskDirect answer, steps and sources
Case studyInspect real executionContext, work, constraints and result
Policy/utilityResolve a durable requirementAccurate terms, contact or tool

Start with customer journeys—not the keyword export

A service business normally needs a useful commercial core before hundreds of articles. Map how a visitor discovers the offer, evaluates fit, verifies capability, handles objections and takes action. A keyword list can reveal demand, but it cannot decide the business journey on its own.

StageUser questionPossible destination
DiscoverCan this solve my problem?Homepage, hub or guide
UnderstandWhat does it involve?Service, methodology or glossary
CompareWhich option fits?Comparison, package or alternatives
VerifyHas this been done credibly?Case study, portfolio or approved testimonial
DecideWhat will happen and cost?Process, pricing, FAQ or contact
ContinueWhat should I do next?Checklist, related guide or maintenance resource

Build a complete URL inventory before drawing a new tree

A crawl reveals reachable URLs, not every URL the organization has created. Combine crawls with CMS or database exports, XML Sitemaps, Search Console, analytics, server logs, backlink exports and business records. Keep live, redirected, removed, private and staging URLs in the inventory until their state is confirmed.

SourceFindsImportant limitation
CrawlerReachable pages, links, status and depthCannot find true orphans
CMS/databasePublished, draft and generated recordsMay include private or obsolete URLs
XML SitemapDeclared preferred URLsA hint, not proof of crawl or index
Search ConsoleGoogle-known and performing URLsSamples and report scope differ
AnalyticsURLs with recorded visitsNo visit does not mean no page
Server logsReal crawler and user requestsRequires retention and parsing
Backlink exportExternally referenced URLsDoes not show internal relevance

Create the page map before the navigation

The page map is a decision register, not a visual sitemap alone. Each proposed URL needs a job, parent, audience, intent, index state, primary conversion, proof requirement and owner. Navigation is then designed from approved journeys instead of deciding the information architecture by accident.

FieldQuestion
Preferred URLWhat stable address owns the page?
Page typeHub, service, guide, case study, listing or utility?
Audience and intentWho needs it and what are they trying to do?
Parent/siblingsWhere does it belong?
Incoming pathWhich approved page introduces it?
Index intentIndex, noindex, private, redirect or retire?
EvidenceWhat makes this page distinct and trustworthy?
Owner/reviewWho maintains it and when?

Decide whether a query deserves a new URL

Search volume alone is not a page-creation rule. Create a separate page when the audience task, required answer, conversion, evidence or product set is meaningfully different. If the same page can satisfy the variation without becoming confusing, strengthen the existing owner URL.

SignalUsually separate URLUsually same URL
IntentDifferent task or decisionSame task with wording variation
OfferDistinct scope, eligibility or conversionMinor feature difference
AudienceMaterially different needs or rulesSame answer with examples
LocationReal local operation and unique evidenceOnly city name changed
LanguageFully localized equivalentOnly navigation translated
FormatTool, calculator or case study has unique jobFAQ can sit within the main page

Use a logical hierarchy without chasing a fixed depth

Important journeys should be direct from relevant hubs and navigation. This does not mean every URL must be one click from the homepage or within a universal three-click rule; Google does not publish such a requirement.

Depth is measured from a chosen entry point and changes with templates. Diagnose needless detours and true isolation rather than forcing every page into a flat list.

LevelTypical roleSEOWithJack example
1Homepage and primary routes/
2Core hubs/services/, /blog/, /portfolio/web-design/
3Specific offer or subject/services/seo-geo/, /hub/seo/
4Detailed guide or case study when useful/internal-linking-for-seo/ or a project page

Separate URL structure from site structure

Google says it generally infers site structure from page linkages rather than reading the folder pattern as a hierarchy. A neat folder can help people and operations, but it does not replace menus, category paths and contextual links.

A flat-looking URL can belong deep in the content system, while a deeply nested URL can be linked prominently. Judge both naming and link relationships.

Choose readable, durable preferred URLs

Use readable words, valid encoding and consistent conventions. Avoid random IDs, session identifiers, unnecessary parameters and fragments that load materially different indexable content. For JavaScript route changes, use real URLs with the History API.

Do not change a useful established URL merely to add a keyword, remove a folder or make a diagram look cleaner. Every URL change creates redirect, internal-link, Canonical, Sitemap, hreflang and measurement work.

Preferred URL standard
  • Represents the page job rather than a temporary template
  • Uses one lowercase and hyphen convention
  • Avoids index.html, tracking parameters and session IDs
  • Uses predictable encoding and parameter separators
  • Has one HTTPS hostname and trailing-path policy
  • Can survive design and CMS changes
  • Has an equivalent redirect only when a move is necessary
  • Appears consistently in links, Canonical, hreflang and Sitemap

Use folders when they improve management and understanding

Folder decisionUseful whenAvoid when
/services/A stable family of offers needs a hubExisting strong URLs would move for cosmetics
/blog/ or /resources/Editorial content needs shared operations and routingThe folder becomes a mixed dumping ground
/locations/Real location pages share governanceEvery city is generated without local value
/ms/ and /zh-cn/Language versions are fully localizedOnly header and footer change
/products/category/Product discovery follows maintained categoriesFilters create endless path combinations

Design primary navigation around durable choices

Primary navigation should expose a small set of lasting user tasks, not mirror every page. Use specific labels, predictable interaction and the same important destinations on mobile. A dropdown or mega menu can support grouping, but it must remain keyboard, touch and crawl accessible.

Google-generated sitelinks are automated. Clear titles, headings, structure and relevant internal anchors can help systems understand important paths, but no navigation design guarantees a particular sitelink.

Navigation release checks
  • Labels describe the destination without insider language
  • Every navigation destination is a normal <a href>
  • Keyboard focus order and Escape behaviour work
  • Touch users do not depend on hover
  • Mobile includes the same important destinations
  • Current and expanded states are understandable
  • No empty headings or duplicate destinations
  • Links resolve directly to preferred URLs

Use breadcrumbs to communicate a real hierarchy

Breadcrumbs help people orient themselves and can expose a clear parent path. They do not need to reproduce browser history, and their hierarchy should agree with the page map.

Breadcrumb structured data can help Google understand the trail, but markup must represent visible or defensible page relationships. Structured data reinforces content; it does not create architecture by itself.

PatternGood useRisk
Home → Services → SEO / GEOStable service hierarchyParent differs across templates
Home → SEO Hub → Internal LinkingLearning routeBreadcrumb points to a non-existent hub
Home → Category → ProductBrowse routeFiltered state is used as permanent parent
Home → Search result → ProductPoor breadcrumb modelReflects temporary visit history

A hub helps an audience choose among maintained subtopics. Contextual links connect prerequisites, explanations, evidence and next steps that do not belong in global navigation. Continue with the internal linking guide for anchors, orphan discovery and edge-level QA.

Do not create an empty hub merely to add a folder level. A hub needs a distinct routing job, clear introduction, curated destinations and ongoing ownership.

Do not rely on site search for discovery

Googlebot generally does not submit searches into a site search box. Pages you want crawled should be reachable through normal links such as categories, pagination, hubs or contextual paths. A Sitemap may supplement discovery when linking every item is difficult, but it does not replace the user journey.

Site-search results are often dynamic, duplicative and low-value as landing pages. Give them an intentional crawl and index policy rather than allowing every query URL to expand the site.

Assign an explicit index state to every page type

StateUse whenTechnical expression
IndexDistinct public page should participate in Search200, crawlable, self-consistent Canonical and internal path
NoindexPublicly accessible but should not appearCrawlable page-level or header noindex
PrivateAccess must be restrictedAuthentication or authorization
Canonical duplicateVariant is useful but substantially duplicate200 plus correct Canonical relationship
RedirectOld URL has an equivalent replacement301/308 to final preferred URL
RetireNo resource or equivalent remainsHonest 404 or 410
Blocked crawlKnown infinite/low-value crawl spaceCarefully scoped robots.txt rule

Make Canonical signals agree

Canonicalization is Google selecting a representative URL among duplicate or highly similar pages. A declared Canonical is a strong signal, not an instruction Google must follow. Internal links, redirects, Sitemaps, HTTPS and hreflang clusters can also influence the selected version.

Use an absolute self-referential Canonical on preferred HTML pages. Do not list duplicate URLs in the Sitemap or point internal links at non-preferred variants.

Canonical consistency gate
  • Preferred page returns 200 and is indexable
  • Canonical is absolute and appears in valid <head>
  • Duplicate variants point to a genuine equivalent
  • Internal links use the preferred protocol, host and path
  • Sitemap contains only preferred indexable URLs
  • Redirected URLs are not Canonical or hreflang targets
  • Localized equivalents self-canonicalize
  • Google-selected Canonical is monitored after release

Control duplicate and near-duplicate templates at the source

CauseArchitecture responseDo not do
Tracking parametersKeep preferred internal URLs and Canonical consistencyAdd campaign parameters to navigation
Print/share variantsConsolidate or noindex based on user needIndex every presentation copy
HTTP/HTTPS or host variantsRedirect to one HTTPS hostServe both as independent sites
CMS archivesKeep only archives with a maintained jobIndex every date, author and tag by default
Location templatesRequire genuine local differenceSwap city names across identical pages
Staging/demoRestrict accessRely only on Canonical to production

Design pagination as part of the architecture

Google crawlers generally do not click Load More or trigger user actions. Paginated component pages need distinct URLs and sequential normal links so deep items remain discoverable. URL fragments are not suitable page identifiers for indexable content.

Each component page normally contains different items and should use an appropriate self-referential Canonical. Google does not use rel="next" and rel="prev" for indexing.

Pagination controls
  • Each component page has a stable URL
  • Next and previous controls are normal anchors
  • Deep items work without interaction or scroll
  • Non-existent page numbers return 404
  • Each component page has an appropriate Canonical
  • Filters and sorting have separate rules
  • Sitemap includes only preferred landing URLs
  • Fresh crawl reaches representative deep items

Treat faceted navigation as a URL-space decision

Filters can be excellent for users while generating near-infinite combinations for crawlers. Decide which combinations, if any, deserve search landing pages before development. If faceted URLs do not need indexing, prevent or constrain their crawl paths with carefully tested controls.

If selected facets are indexable, enforce a consistent parameter or path order, prevent duplicate filters, return 404 for empty or nonsensical combinations and provide unique value. Canonical and nofollow may reduce crawling over time but are weaker controls than preventing unwanted crawl paths at the design level.

Facet typeLikely policyReason
Sort orderUsually non-indexableSame items, presentation changes
Session/view stateNever an index landing pageNo stable search value
Popular product attributeSelected indexable landing pageDistinct demand and useful inventory
Multiple additive filtersUsually constrainedCombinatorial URL growth
Zero-result combination404No resource exists
Internal site-search queryUsually non-indexableUncontrolled and duplicative

Govern categories, tags, authors, dates and search archives

An archive deserves indexing only when it has a stable audience job, enough useful items, crawlable pagination, clear ownership and distinct value beyond a list. The existence of a CMS taxonomy does not create search demand.

Consolidate synonymous taxonomies, remove one-item archives from important journeys and prevent editors from creating uncontrolled tags. Author pages can be useful when authorship matters and the page offers real profile value; date archives rarely need to be search landing pages for an evergreen site.

Build JavaScript and SPA routes as independent pages

Every indexable view needs a resolvable URL, appropriate server response, real <a href> paths, rendered primary content, unique metadata and consistent Canonical. Use the History API instead of hash fragments for different pages.

Test deep routes directly, not only after entering through the homepage. A client-side view that returns an app shell, wrong status or missing content to a direct request is an architecture defect.

Preserve mobile discovery parity

Google indexes with the mobile version. Important content and internal links present on desktop should remain available in the mobile rendered version. Compact menus and accordions can be used when they expose the same meaningful destinations.

Architecture QA must include touch, keyboard, visible focus, accessible names, responsive tables and links that do not appear only on hover.

Plan multilingual architecture before localization

Use distinct, stable URLs for each language. SEOWithJack uses English on the primary route, Bahasa Melayu under /ms/ and Simplified Chinese under /zh-cn/. Each genuine translation should self-canonicalize and participate in a reciprocal hreflang cluster.

Translate intent, navigation, anchors, metadata and conversion journeys—not only the article body. Never Canonicalize a complete Malay or Chinese equivalent to English, and avoid automatic redirection based only on IP or assumed language. See the multilingual SEO guide.

Locale architecture gate
  • Equivalent pages have distinct stable URLs
  • HTML language and visible copy agree
  • Each version has a self-referential Canonical
  • Hreflang references are reciprocal and valid
  • Language switcher uses normal links to equivalents
  • Missing translations have an honest fallback
  • Navigation and CTAs stay in the selected language
  • Sitemap and internal links use preferred locale URLs

Use structured data to reinforce—not invent—structure

Breadcrumb, Organization, Article, Product and other eligible structured data can clarify entities and page roles, but markup must match visible content and supported features. It does not compensate for missing category links, thin pages or contradictory Canonicals.

Apply schema by page type through maintained templates, validate generated output and remove properties that are not true.

Choose an architecture pattern that fits the site

Site typeDurable coreKey risk
Small service businessHomepage → services → proof/process/contactBlog grows before commercial pages
Content publisherTopic hubs → guides → related tasksTags and dates fragment coverage
EcommerceCategories → subcategories → productsFacets and pagination create crawl space
Marketplace/directoryBrowse routes → entity/profile pagesEmpty combinations and thin profiles
SaaS/productUse cases/features/resources/docsMarketing and documentation duplicate intent
Multilingual businessEquivalent locale paths and local evidencePartial translation and broken hreflang

Find and decide true orphan pages

A true orphan has no discoverable incoming internal path in the tested site graph. Compare the crawl with Sitemaps, CMS, analytics, Search Console, logs and backlink records. Then decide whether the page belongs in the architecture before adding links.

Important unique pages need an approved parent and contextual entry. Duplicate pages may be consolidated, private pages protected, obsolete pages retired and staging artifacts removed. More links are not the correct answer for every orphan.

Use Sitemaps as a declaration, not the architecture

A Sitemap tells Google which preferred URLs you want considered. Submission is a hint and does not guarantee crawling, indexing or ranking. Include only Canonical, indexable preferred URLs and use accurate lastmod values only for significant updates.

A Sitemap can supplement discovery for large inventories, but normal category and contextual paths remain necessary for users and help express page relationships.

Treat crawl budget in proportion to site scale

Most small sites do not need an elaborate crawl-budget project. The priority is removing broken journeys, duplicate URL generation and accidental index spaces. Larger or fast-changing sites should use server logs, Crawl Stats and segmented inventories to understand demand and capacity.

Facets, session IDs, duplicate content, soft errors, hacked URLs and infinite spaces can consume resources. Fix generation and linking rules before chasing a generic crawl score.

Protect architecture during redesigns and migrations

Freeze the approved page map before launch, inventory old URLs and assign every changed URL a status. Use server-side permanent redirects only to genuinely equivalent destinations, update all internal signals to final URLs and avoid redirect chains.

Compare old, staging and production crawls. Monitor important templates, 404s, Google-selected Canonicals, indexed cohorts and user journeys after launch. Keep redirects long enough for users and systems; do not mass-redirect removed unrelated pages to the homepage.

StageControlEvidence
InventoryCapture all old URLs and relationshipsCrawl, CMS, Sitemap, logs and backlinks
DecisionKeep, improve, merge, move, retire or restrictApproved mapping and owner
StagingUse final internal links and metadataNo avoidable internal redirects
LaunchRelease redirects, Sitemap and robots togetherStatus and route tests
After launchMonitor cohorts and unexpected CanonicalsSearch Console, logs, analytics and recrawls

Make architecture change-controlled

Architecture decays when new pages, campaigns, tags and filters can be created without ownership. Establish approved page types, URL rules, navigation eligibility, index defaults, redirect requirements and deletion procedures in the CMS or release process.

A page request should explain the user task, existing owner URL, expected evidence, parent path, index intent and maintenance owner before development starts.

ChangeApproval questionRelease evidence
New pageWhy can no existing URL own this task?Page brief and incoming path
New taxonomyWill it remain useful with enough items?Eligibility and empty-state rules
New filterCan it generate crawlable combinations?Parameter and index policy
URL changeIs the benefit worth migration risk?Redirect and signal update map
Page removalIs there an equivalent destination?Status, link and Sitemap update
Navigation changeWhich durable journey improves?Desktop, mobile and keyboard QA

Measure architecture by outcomes, not a single score

LayerMeasureQuestion
IntegrityValid preferred URLs and signalsIs the structure technically consistent?
DiscoveryReachable indexable pages and crawl evidenceCan systems find the pages?
SelectionGoogle-selected Canonical and indexed cohortAre intended versions participating?
JourneyNavigation and contextual-link usageCan people reach useful next steps?
SearchQuery-to-owner-page alignmentIs the correct page appearing?
BusinessQualified actions by landing pathDoes the structure support decisions?
MaintenanceNew orphans, duplicates and redirect debtIs architecture staying healthy?

Diagnose patterns before changing the tree

PatternChecksDo not assume
Important page not crawledIncoming href, robots, status, Sitemap and logsSubmitting Sitemap guarantees discovery
Wrong URL appearsIntent ownership, Canonical, links and duplicationFolder depth alone caused it
Many crawled-not-indexed pagesTemplate value, duplication, facets and index intentAll need more internal links
Deep products missedCategory pagination and Load More pathsSearch box is sufficient
Language version absentBody localization, Canonical and reciprocal hreflangGoogle will infer the language route
Traffic falls after migrationRedirects, signals, parity, demand and trackingOne architecture metric explains everything

Architecture myths to retire

  • Google does not publish a universal three-click rule or ideal number of hierarchy levels.
  • A flat structure does not automatically rank better.
  • URL folders alone do not tell Google the complete site hierarchy; page linkages matter.
  • Adding keywords to every slug is not an architecture strategy.
  • A Sitemap supplements discovery but does not replace normal internal paths or guarantee indexing.
  • Breadcrumb markup does not create a hierarchy that the website itself does not support.
  • More categories, tags, filters and location pages do not create more authority.
  • Canonical is a signal, not a directive that can safely solve every duplicate URL at scale.
  • Robots.txt prevents crawling; it is not a reliable way to remove an already known URL from Search.
  • Every orphan page does not deserve a link.
  • Redirecting every removed page to the homepage is not a valid consolidation strategy.
  • No architecture can guarantee crawling, indexing, sitelinks, rankings, traffic or conversions.

A repeatable architecture workflow

From inventory to governed release
  1. Define audiences, offers, tasks and evidence requirements.
  2. Combine crawl, CMS, Sitemap, Search Console, analytics, logs and backlink inventories.
  3. Assign every existing URL a job and index state.
  4. Group demand by intent and identify one owner URL per task.
  5. Approve the commercial core, hubs and support pages.
  6. Map parent, sibling, incoming and conversion paths.
  7. Set preferred URL, Canonical, hreflang, pagination, facet and archive rules.
  8. Design accessible desktop and mobile navigation from approved journeys.
  9. Prepare redirects before changing or consolidating any URL.
  10. Build and crawl staging from normal entry points.
  11. Test direct routes, rendered mobile pages, status and signals.
  12. Release through change control and monitor page-type cohorts.

Frequently asked questions

How many levels should a site have?

There is no universal number. Keep priority journeys direct and add depth only when real information or product relationships require it.

Does a flat architecture rank better?

Not automatically. A coherent hierarchy with relevant links is more useful than forcing every URL into one level.

Does Google use URL folders to understand hierarchy?

Readable folders help users and operations, but Google says link relationships are a major way it understands site structure.

Should blog posts use a /blog/ folder?

Either pattern can work. Choose a durable convention and do not move established URLs only for cosmetic consistency.

No. It can expose preferred URLs as a hint, but it does not create a browse journey or guarantee crawl and index.

Should tag and filter pages be indexed?

Only selected pages with a stable audience job, distinct value and ongoing ownership. Do not index every generated combination.

Are breadcrumbs required?

No, but a truthful visible trail can improve orientation. Structured data must match the actual page relationship.

What should happen to an old URL?

Keep it when still suitable; otherwise redirect to a genuinely equivalent replacement, or return 404/410 when no replacement exists.

How should multilingual pages be structured?

Use stable locale URLs, fully localize the experience, self-canonicalize each version and connect equivalents with reciprocal hreflang.

When should architecture be audited?

Review it before redesigns, migrations, new taxonomies and major content expansion, then monitor recurring orphan, duplicate and redirect patterns.

Official references

Need a practical next step?Turn the URL inventory into a governed page map.

Share the current Sitemap, services and planned content. Jack can identify page ownership, unnecessary URL spaces, missing journeys and migration risks before development scales them.

Discuss site architecture on WhatsApp

Jack Lee

Jack Lee

Building Search Visibility with SEO, GEO & AI-Assisted Websites through practical projects and experiments.