SEO site architecture is the operating model behind the website: which pages exist, what job each URL owns, how people and crawlers reach them, which variants may be indexed and how the structure remains coherent as the business changes. It is broader than a menu and cannot be reduced to a folder pattern or a fixed click-depth score.
Create one durable preferred page for each necessary user task, then connect it through crawlable journeys, consistent index signals and an owner who can keep it accurate.
Define architecture as a set of controlled decisions
| Layer | Decision | Required output |
|---|---|---|
| Audience | Who must complete which task? | Journey and evidence map |
| Page system | Which URL owns each task? | Approved page map |
| Hierarchy | How are pages grouped? | Parent, child and sibling relationships |
| Discovery | How are pages reached? | Navigation and contextual link paths |
| Index control | Which versions may appear in Search? | Canonical, noindex, robots and status rules |
| Operations | Who approves change? | Owner, trigger and QA record |
Give every page one primary job
A page may support several secondary actions, but it needs one primary audience, intent and outcome. When two URLs own the same job, they compete for maintenance, links, measurement and user attention even if Google does not treat them as a ranking problem.
Use search intent and first-party customer questions to define the job before writing a slug.
| Page job | Primary outcome | Typical evidence |
|---|---|---|
| Homepage | Orient and route | Clear positioning and major choices |
| Hub | Choose a path within a subject | Curated child resources |
| Service/product | Evaluate an offer | Scope, fit, proof and next step |
| Guide | Complete a learning or operating task | Direct answer, steps and sources |
| Case study | Inspect real execution | Context, work, constraints and result |
| Policy/utility | Resolve a durable requirement | Accurate terms, contact or tool |
Start with customer journeys—not the keyword export
A service business normally needs a useful commercial core before hundreds of articles. Map how a visitor discovers the offer, evaluates fit, verifies capability, handles objections and takes action. A keyword list can reveal demand, but it cannot decide the business journey on its own.
| Stage | User question | Possible destination |
|---|---|---|
| Discover | Can this solve my problem? | Homepage, hub or guide |
| Understand | What does it involve? | Service, methodology or glossary |
| Compare | Which option fits? | Comparison, package or alternatives |
| Verify | Has this been done credibly? | Case study, portfolio or approved testimonial |
| Decide | What will happen and cost? | Process, pricing, FAQ or contact |
| Continue | What should I do next? | Checklist, related guide or maintenance resource |
Build a complete URL inventory before drawing a new tree
A crawl reveals reachable URLs, not every URL the organization has created. Combine crawls with CMS or database exports, XML Sitemaps, Search Console, analytics, server logs, backlink exports and business records. Keep live, redirected, removed, private and staging URLs in the inventory until their state is confirmed.
| Source | Finds | Important limitation |
|---|---|---|
| Crawler | Reachable pages, links, status and depth | Cannot find true orphans |
| CMS/database | Published, draft and generated records | May include private or obsolete URLs |
| XML Sitemap | Declared preferred URLs | A hint, not proof of crawl or index |
| Search Console | Google-known and performing URLs | Samples and report scope differ |
| Analytics | URLs with recorded visits | No visit does not mean no page |
| Server logs | Real crawler and user requests | Requires retention and parsing |
| Backlink export | Externally referenced URLs | Does not show internal relevance |
Create the page map before the navigation
The page map is a decision register, not a visual sitemap alone. Each proposed URL needs a job, parent, audience, intent, index state, primary conversion, proof requirement and owner. Navigation is then designed from approved journeys instead of deciding the information architecture by accident.
| Field | Question |
|---|---|
| Preferred URL | What stable address owns the page? |
| Page type | Hub, service, guide, case study, listing or utility? |
| Audience and intent | Who needs it and what are they trying to do? |
| Parent/siblings | Where does it belong? |
| Incoming path | Which approved page introduces it? |
| Index intent | Index, noindex, private, redirect or retire? |
| Evidence | What makes this page distinct and trustworthy? |
| Owner/review | Who maintains it and when? |
Decide whether a query deserves a new URL
Search volume alone is not a page-creation rule. Create a separate page when the audience task, required answer, conversion, evidence or product set is meaningfully different. If the same page can satisfy the variation without becoming confusing, strengthen the existing owner URL.
| Signal | Usually separate URL | Usually same URL |
|---|---|---|
| Intent | Different task or decision | Same task with wording variation |
| Offer | Distinct scope, eligibility or conversion | Minor feature difference |
| Audience | Materially different needs or rules | Same answer with examples |
| Location | Real local operation and unique evidence | Only city name changed |
| Language | Fully localized equivalent | Only navigation translated |
| Format | Tool, calculator or case study has unique job | FAQ can sit within the main page |
Use a logical hierarchy without chasing a fixed depth
Important journeys should be direct from relevant hubs and navigation. This does not mean every URL must be one click from the homepage or within a universal three-click rule; Google does not publish such a requirement.
Depth is measured from a chosen entry point and changes with templates. Diagnose needless detours and true isolation rather than forcing every page into a flat list.
| Level | Typical role | SEOWithJack example |
|---|---|---|
| 1 | Homepage and primary routes | / |
| 2 | Core hubs | /services/, /blog/, /portfolio/web-design/ |
| 3 | Specific offer or subject | /services/seo-geo/, /hub/seo/ |
| 4 | Detailed guide or case study when useful | /internal-linking-for-seo/ or a project page |
Separate URL structure from site structure
Google says it generally infers site structure from page linkages rather than reading the folder pattern as a hierarchy. A neat folder can help people and operations, but it does not replace menus, category paths and contextual links.
A flat-looking URL can belong deep in the content system, while a deeply nested URL can be linked prominently. Judge both naming and link relationships.
Choose readable, durable preferred URLs
Use readable words, valid encoding and consistent conventions. Avoid random IDs, session identifiers, unnecessary parameters and fragments that load materially different indexable content. For JavaScript route changes, use real URLs with the History API.
Do not change a useful established URL merely to add a keyword, remove a folder or make a diagram look cleaner. Every URL change creates redirect, internal-link, Canonical, Sitemap, hreflang and measurement work.
- Represents the page job rather than a temporary template
- Uses one lowercase and hyphen convention
- Avoids
index.html, tracking parameters and session IDs - Uses predictable encoding and parameter separators
- Has one HTTPS hostname and trailing-path policy
- Can survive design and CMS changes
- Has an equivalent redirect only when a move is necessary
- Appears consistently in links, Canonical, hreflang and Sitemap
Use folders when they improve management and understanding
| Folder decision | Useful when | Avoid when |
|---|---|---|
| /services/ | A stable family of offers needs a hub | Existing strong URLs would move for cosmetics |
| /blog/ or /resources/ | Editorial content needs shared operations and routing | The folder becomes a mixed dumping ground |
| /locations/ | Real location pages share governance | Every city is generated without local value |
| /ms/ and /zh-cn/ | Language versions are fully localized | Only header and footer change |
| /products/category/ | Product discovery follows maintained categories | Filters create endless path combinations |
Design primary navigation around durable choices
Primary navigation should expose a small set of lasting user tasks, not mirror every page. Use specific labels, predictable interaction and the same important destinations on mobile. A dropdown or mega menu can support grouping, but it must remain keyboard, touch and crawl accessible.
Google-generated sitelinks are automated. Clear titles, headings, structure and relevant internal anchors can help systems understand important paths, but no navigation design guarantees a particular sitelink.
- Labels describe the destination without insider language
- Every navigation destination is a normal
<a href> - Keyboard focus order and Escape behaviour work
- Touch users do not depend on hover
- Mobile includes the same important destinations
- Current and expanded states are understandable
- No empty headings or duplicate destinations
- Links resolve directly to preferred URLs
Use breadcrumbs to communicate a real hierarchy
Breadcrumbs help people orient themselves and can expose a clear parent path. They do not need to reproduce browser history, and their hierarchy should agree with the page map.
Breadcrumb structured data can help Google understand the trail, but markup must represent visible or defensible page relationships. Structured data reinforces content; it does not create architecture by itself.
| Pattern | Good use | Risk |
|---|---|---|
| Home → Services → SEO / GEO | Stable service hierarchy | Parent differs across templates |
| Home → SEO Hub → Internal Linking | Learning route | Breadcrumb points to a non-existent hub |
| Home → Category → Product | Browse route | Filtered state is used as permanent parent |
| Home → Search result → Product | Poor breadcrumb model | Reflects temporary visit history |
Use hubs and contextual links for the relationships menus cannot show
A hub helps an audience choose among maintained subtopics. Contextual links connect prerequisites, explanations, evidence and next steps that do not belong in global navigation. Continue with the internal linking guide for anchors, orphan discovery and edge-level QA.
Do not create an empty hub merely to add a folder level. A hub needs a distinct routing job, clear introduction, curated destinations and ongoing ownership.
Do not rely on site search for discovery
Googlebot generally does not submit searches into a site search box. Pages you want crawled should be reachable through normal links such as categories, pagination, hubs or contextual paths. A Sitemap may supplement discovery when linking every item is difficult, but it does not replace the user journey.
Site-search results are often dynamic, duplicative and low-value as landing pages. Give them an intentional crawl and index policy rather than allowing every query URL to expand the site.
Assign an explicit index state to every page type
| State | Use when | Technical expression |
|---|---|---|
| Index | Distinct public page should participate in Search | 200, crawlable, self-consistent Canonical and internal path |
| Noindex | Publicly accessible but should not appear | Crawlable page-level or header noindex |
| Private | Access must be restricted | Authentication or authorization |
| Canonical duplicate | Variant is useful but substantially duplicate | 200 plus correct Canonical relationship |
| Redirect | Old URL has an equivalent replacement | 301/308 to final preferred URL |
| Retire | No resource or equivalent remains | Honest 404 or 410 |
| Blocked crawl | Known infinite/low-value crawl space | Carefully scoped robots.txt rule |
Make Canonical signals agree
Canonicalization is Google selecting a representative URL among duplicate or highly similar pages. A declared Canonical is a strong signal, not an instruction Google must follow. Internal links, redirects, Sitemaps, HTTPS and hreflang clusters can also influence the selected version.
Use an absolute self-referential Canonical on preferred HTML pages. Do not list duplicate URLs in the Sitemap or point internal links at non-preferred variants.
- Preferred page returns 200 and is indexable
- Canonical is absolute and appears in valid
<head> - Duplicate variants point to a genuine equivalent
- Internal links use the preferred protocol, host and path
- Sitemap contains only preferred indexable URLs
- Redirected URLs are not Canonical or hreflang targets
- Localized equivalents self-canonicalize
- Google-selected Canonical is monitored after release
Control duplicate and near-duplicate templates at the source
| Cause | Architecture response | Do not do |
|---|---|---|
| Tracking parameters | Keep preferred internal URLs and Canonical consistency | Add campaign parameters to navigation |
| Print/share variants | Consolidate or noindex based on user need | Index every presentation copy |
| HTTP/HTTPS or host variants | Redirect to one HTTPS host | Serve both as independent sites |
| CMS archives | Keep only archives with a maintained job | Index every date, author and tag by default |
| Location templates | Require genuine local difference | Swap city names across identical pages |
| Staging/demo | Restrict access | Rely only on Canonical to production |
Design pagination as part of the architecture
Google crawlers generally do not click Load More or trigger user actions. Paginated component pages need distinct URLs and sequential normal links so deep items remain discoverable. URL fragments are not suitable page identifiers for indexable content.
Each component page normally contains different items and should use an appropriate self-referential Canonical. Google does not use rel="next" and rel="prev" for indexing.
- Each component page has a stable URL
- Next and previous controls are normal anchors
- Deep items work without interaction or scroll
- Non-existent page numbers return 404
- Each component page has an appropriate Canonical
- Filters and sorting have separate rules
- Sitemap includes only preferred landing URLs
- Fresh crawl reaches representative deep items
Treat faceted navigation as a URL-space decision
Filters can be excellent for users while generating near-infinite combinations for crawlers. Decide which combinations, if any, deserve search landing pages before development. If faceted URLs do not need indexing, prevent or constrain their crawl paths with carefully tested controls.
If selected facets are indexable, enforce a consistent parameter or path order, prevent duplicate filters, return 404 for empty or nonsensical combinations and provide unique value. Canonical and nofollow may reduce crawling over time but are weaker controls than preventing unwanted crawl paths at the design level.
| Facet type | Likely policy | Reason |
|---|---|---|
| Sort order | Usually non-indexable | Same items, presentation changes |
| Session/view state | Never an index landing page | No stable search value |
| Popular product attribute | Selected indexable landing page | Distinct demand and useful inventory |
| Multiple additive filters | Usually constrained | Combinatorial URL growth |
| Zero-result combination | 404 | No resource exists |
| Internal site-search query | Usually non-indexable | Uncontrolled and duplicative |
Govern categories, tags, authors, dates and search archives
An archive deserves indexing only when it has a stable audience job, enough useful items, crawlable pagination, clear ownership and distinct value beyond a list. The existence of a CMS taxonomy does not create search demand.
Consolidate synonymous taxonomies, remove one-item archives from important journeys and prevent editors from creating uncontrolled tags. Author pages can be useful when authorship matters and the page offers real profile value; date archives rarely need to be search landing pages for an evergreen site.
Build JavaScript and SPA routes as independent pages
Every indexable view needs a resolvable URL, appropriate server response, real <a href> paths, rendered primary content, unique metadata and consistent Canonical. Use the History API instead of hash fragments for different pages.
Test deep routes directly, not only after entering through the homepage. A client-side view that returns an app shell, wrong status or missing content to a direct request is an architecture defect.
Preserve mobile discovery parity
Google indexes with the mobile version. Important content and internal links present on desktop should remain available in the mobile rendered version. Compact menus and accordions can be used when they expose the same meaningful destinations.
Architecture QA must include touch, keyboard, visible focus, accessible names, responsive tables and links that do not appear only on hover.
Plan multilingual architecture before localization
Use distinct, stable URLs for each language. SEOWithJack uses English on the primary route, Bahasa Melayu under /ms/ and Simplified Chinese under /zh-cn/. Each genuine translation should self-canonicalize and participate in a reciprocal hreflang cluster.
Translate intent, navigation, anchors, metadata and conversion journeys—not only the article body. Never Canonicalize a complete Malay or Chinese equivalent to English, and avoid automatic redirection based only on IP or assumed language. See the multilingual SEO guide.
- Equivalent pages have distinct stable URLs
- HTML language and visible copy agree
- Each version has a self-referential Canonical
- Hreflang references are reciprocal and valid
- Language switcher uses normal links to equivalents
- Missing translations have an honest fallback
- Navigation and CTAs stay in the selected language
- Sitemap and internal links use preferred locale URLs
Use structured data to reinforce—not invent—structure
Breadcrumb, Organization, Article, Product and other eligible structured data can clarify entities and page roles, but markup must match visible content and supported features. It does not compensate for missing category links, thin pages or contradictory Canonicals.
Apply schema by page type through maintained templates, validate generated output and remove properties that are not true.
Choose an architecture pattern that fits the site
| Site type | Durable core | Key risk |
|---|---|---|
| Small service business | Homepage → services → proof/process/contact | Blog grows before commercial pages |
| Content publisher | Topic hubs → guides → related tasks | Tags and dates fragment coverage |
| Ecommerce | Categories → subcategories → products | Facets and pagination create crawl space |
| Marketplace/directory | Browse routes → entity/profile pages | Empty combinations and thin profiles |
| SaaS/product | Use cases/features/resources/docs | Marketing and documentation duplicate intent |
| Multilingual business | Equivalent locale paths and local evidence | Partial translation and broken hreflang |
Find and decide true orphan pages
A true orphan has no discoverable incoming internal path in the tested site graph. Compare the crawl with Sitemaps, CMS, analytics, Search Console, logs and backlink records. Then decide whether the page belongs in the architecture before adding links.
Important unique pages need an approved parent and contextual entry. Duplicate pages may be consolidated, private pages protected, obsolete pages retired and staging artifacts removed. More links are not the correct answer for every orphan.
Use Sitemaps as a declaration, not the architecture
A Sitemap tells Google which preferred URLs you want considered. Submission is a hint and does not guarantee crawling, indexing or ranking. Include only Canonical, indexable preferred URLs and use accurate lastmod values only for significant updates.
A Sitemap can supplement discovery for large inventories, but normal category and contextual paths remain necessary for users and help express page relationships.
Treat crawl budget in proportion to site scale
Most small sites do not need an elaborate crawl-budget project. The priority is removing broken journeys, duplicate URL generation and accidental index spaces. Larger or fast-changing sites should use server logs, Crawl Stats and segmented inventories to understand demand and capacity.
Facets, session IDs, duplicate content, soft errors, hacked URLs and infinite spaces can consume resources. Fix generation and linking rules before chasing a generic crawl score.
Protect architecture during redesigns and migrations
Freeze the approved page map before launch, inventory old URLs and assign every changed URL a status. Use server-side permanent redirects only to genuinely equivalent destinations, update all internal signals to final URLs and avoid redirect chains.
Compare old, staging and production crawls. Monitor important templates, 404s, Google-selected Canonicals, indexed cohorts and user journeys after launch. Keep redirects long enough for users and systems; do not mass-redirect removed unrelated pages to the homepage.
| Stage | Control | Evidence |
|---|---|---|
| Inventory | Capture all old URLs and relationships | Crawl, CMS, Sitemap, logs and backlinks |
| Decision | Keep, improve, merge, move, retire or restrict | Approved mapping and owner |
| Staging | Use final internal links and metadata | No avoidable internal redirects |
| Launch | Release redirects, Sitemap and robots together | Status and route tests |
| After launch | Monitor cohorts and unexpected Canonicals | Search Console, logs, analytics and recrawls |
Make architecture change-controlled
Architecture decays when new pages, campaigns, tags and filters can be created without ownership. Establish approved page types, URL rules, navigation eligibility, index defaults, redirect requirements and deletion procedures in the CMS or release process.
A page request should explain the user task, existing owner URL, expected evidence, parent path, index intent and maintenance owner before development starts.
| Change | Approval question | Release evidence |
|---|---|---|
| New page | Why can no existing URL own this task? | Page brief and incoming path |
| New taxonomy | Will it remain useful with enough items? | Eligibility and empty-state rules |
| New filter | Can it generate crawlable combinations? | Parameter and index policy |
| URL change | Is the benefit worth migration risk? | Redirect and signal update map |
| Page removal | Is there an equivalent destination? | Status, link and Sitemap update |
| Navigation change | Which durable journey improves? | Desktop, mobile and keyboard QA |
Measure architecture by outcomes, not a single score
| Layer | Measure | Question |
|---|---|---|
| Integrity | Valid preferred URLs and signals | Is the structure technically consistent? |
| Discovery | Reachable indexable pages and crawl evidence | Can systems find the pages? |
| Selection | Google-selected Canonical and indexed cohort | Are intended versions participating? |
| Journey | Navigation and contextual-link usage | Can people reach useful next steps? |
| Search | Query-to-owner-page alignment | Is the correct page appearing? |
| Business | Qualified actions by landing path | Does the structure support decisions? |
| Maintenance | New orphans, duplicates and redirect debt | Is architecture staying healthy? |
Diagnose patterns before changing the tree
| Pattern | Checks | Do not assume |
|---|---|---|
| Important page not crawled | Incoming href, robots, status, Sitemap and logs | Submitting Sitemap guarantees discovery |
| Wrong URL appears | Intent ownership, Canonical, links and duplication | Folder depth alone caused it |
| Many crawled-not-indexed pages | Template value, duplication, facets and index intent | All need more internal links |
| Deep products missed | Category pagination and Load More paths | Search box is sufficient |
| Language version absent | Body localization, Canonical and reciprocal hreflang | Google will infer the language route |
| Traffic falls after migration | Redirects, signals, parity, demand and tracking | One architecture metric explains everything |
Architecture myths to retire
- Google does not publish a universal three-click rule or ideal number of hierarchy levels.
- A flat structure does not automatically rank better.
- URL folders alone do not tell Google the complete site hierarchy; page linkages matter.
- Adding keywords to every slug is not an architecture strategy.
- A Sitemap supplements discovery but does not replace normal internal paths or guarantee indexing.
- Breadcrumb markup does not create a hierarchy that the website itself does not support.
- More categories, tags, filters and location pages do not create more authority.
- Canonical is a signal, not a directive that can safely solve every duplicate URL at scale.
- Robots.txt prevents crawling; it is not a reliable way to remove an already known URL from Search.
- Every orphan page does not deserve a link.
- Redirecting every removed page to the homepage is not a valid consolidation strategy.
- No architecture can guarantee crawling, indexing, sitelinks, rankings, traffic or conversions.
A repeatable architecture workflow
- Define audiences, offers, tasks and evidence requirements.
- Combine crawl, CMS, Sitemap, Search Console, analytics, logs and backlink inventories.
- Assign every existing URL a job and index state.
- Group demand by intent and identify one owner URL per task.
- Approve the commercial core, hubs and support pages.
- Map parent, sibling, incoming and conversion paths.
- Set preferred URL, Canonical, hreflang, pagination, facet and archive rules.
- Design accessible desktop and mobile navigation from approved journeys.
- Prepare redirects before changing or consolidating any URL.
- Build and crawl staging from normal entry points.
- Test direct routes, rendered mobile pages, status and signals.
- Release through change control and monitor page-type cohorts.
Frequently asked questions
How many levels should a site have?
There is no universal number. Keep priority journeys direct and add depth only when real information or product relationships require it.
Does a flat architecture rank better?
Not automatically. A coherent hierarchy with relevant links is more useful than forcing every URL into one level.
Does Google use URL folders to understand hierarchy?
Readable folders help users and operations, but Google says link relationships are a major way it understands site structure.
Should blog posts use a /blog/ folder?
Either pattern can work. Choose a durable convention and do not move established URLs only for cosmetic consistency.
Can a Sitemap replace category links?
No. It can expose preferred URLs as a hint, but it does not create a browse journey or guarantee crawl and index.
Should tag and filter pages be indexed?
Only selected pages with a stable audience job, distinct value and ongoing ownership. Do not index every generated combination.
Are breadcrumbs required?
No, but a truthful visible trail can improve orientation. Structured data must match the actual page relationship.
What should happen to an old URL?
Keep it when still suitable; otherwise redirect to a genuinely equivalent replacement, or return 404/410 when no replacement exists.
How should multilingual pages be structured?
Use stable locale URLs, fully localize the experience, self-canonicalize each version and connect equivalents with reciprocal hreflang.
When should architecture be audited?
Review it before redesigns, migrations, new taxonomies and major content expansion, then monitor recurring orphan, duplicate and redirect patterns.
Official references
- Google SEO Starter Guide: organize your site
- Google Search Essentials
- Google URL structure best practices
- How Google Search discovers pages
- Google: help Google understand your ecommerce site structure
- Google link best practices
- Google breadcrumb structured data
- Google sitelinks guidance
- Google canonicalization methods
- Google: build and submit a Sitemap
- Google: managing faceted navigation crawling
- Google pagination and incremental loading
- Google JavaScript SEO basics
- Google mobile-first indexing best practices
- Google localized versions and hreflang
- Google site moves and migrations
- Google redirects and Search
- Google robots.txt introduction
- Google HTTP status codes and network errors
- Google crawl budget management
- W3C WAI: page structure and navigation
Share the current Sitemap, services and planned content. Jack can identify page ownership, unnecessary URL spaces, missing journeys and migration risks before development scales them.



How Google Ranking Evolved: From PageRank to Modern Search SystemsSeptember 2, 2026
Google AI Content and SEO: What Is Allowed, What Is Spam, and How to Publish SafelySeptember 2, 2026
Google Florida Update (2003): What We Know, What Remains TheorySeptember 2, 2026