XML Sitemap ialah discovery dan inventory hint: ia memberitahu search engine tentang preferred URL dan file yang dianggap penting. robots.txt ialah access protocol: ia memberitahu compliant crawler URL path yang boleh diminta. Kedua-duanya tidak guarantee indexing atau ranking dan tidak melindungi private information.
Discovery: Sitemap bersama crawlable link. Crawl access: robots.txt. Index exclusion: crawlable noindex atau real removal response. Duplicate consolidation: Canonical atau redirect. Confidentiality: authentication dan authorization.
Lima control ini tidak boleh ditukar ganti
| Outcome diperlukan | Primary control | Evidence berjaya | Jangan ganti dengan |
|---|---|---|---|
| Bantu search engine discover preferred URL | Crawlable internal link dan clean Sitemap | URL linked, listed dan fetchable | Repeated URL Inspection request |
| Halang compliant bot meminta sesuatu path | robots.txt rule untuk crawler dan host itu | Rule match tested URL | noindex yang masih membenarkan crawl |
| Keluarkan accessible page daripada Google Cari | robots meta atau X-Robots-Tag noindex | Google boleh fetch dan baca directive | robots.txt Disallow |
| Remove retired public URL secara permanent | 404/410 atau 301/308 ke true replacement | Final HTTP response sepadan dengan keputusan | Buang daripada Sitemap sahaja |
| Consolidate equivalent duplicate URL | Consistent Canonical signal atau permanent redirect | Preferred URL digunakan dalam link dan Sitemap | Block duplicate dalam robots.txt |
| Protect private atau confidential content | Authentication dan server-side authorization | Unauthorized request tidak boleh retrieve content | robots.txt, noindex atau unlinked URL |
Adakah website memerlukan Sitemap?
Google menyatakan small site sekitar 500 page atau kurang mungkin tidak memerlukan Sitemap jika semua important page boleh dicapai melalui normal link. Namun Sitemap masih berguna untuk small business ketika launch, migration, multilingual expansion dan inventory monitoring.
Sitemap melengkapkan site architecture; ia tidak membaiki orphan page. Gunakan panduan site architecture dan panduan internal linking untuk real discovery path.
| Situasi | Nilai Sitemap | Companion lebih penting |
|---|---|---|
| New domain dengan sedikit external link | Tinggi: tunjuk preferred launch inventory | Link daripada homepage, hub dan relevant external profile |
| Small service site yang comprehensively linked | Useful tetapi tidak essential untuk discovery | Clear navigation dan direct internal link |
| Large ecommerce, publishing atau marketplace | Tinggi: partition inventory besar dan berubah | Faceted-navigation control, log dan stable URL rule |
| Multilingual website | Tinggi jika Sitemap hreflang ialah chosen implementation | Self-Canonical locale URL dan reciprocal alternate |
| Image, video atau news-led site | Tinggi jika specialized file sukar ditemukan | Accessible media, valid landing page dan feature requirement |
| Website migration | Tinggi: tunjuk final preferred URL set | One-to-one redirect, preserved content dan production crawl QA |
Pilih Sitemap format yang boleh dimaintain
| Format | Sesuai untuk | Capability dan limit |
|---|---|---|
| XML URL set | Kebanyakan site dan CMS | Support URL, lastmod serta Google extension untuk locale, image, video dan news |
| Sitemap index | Multiple XML file atau large inventory | Reference child Sitemap; mengandungi Sitemap location, bukan page URL |
| RSS atau Atom | Publisher dengan reliable feed | Useful untuk recent URL; bukan complete historical inventory |
| Text Sitemap | Simple page-only inventory | Satu absolute URL setiap line; tiada lastmod atau extension |
| HTML Sitemap | Human navigation dan accessibility | Normal webpage, bukan pengganti protocol-compliant Sitemap submission |
Bina Sitemap daripada canonical inventory
Generator paling selamat menggunakan source of truth yang sama dengan publishing system. URL hanya masuk Sitemap apabila intended state ialah public, indexable dan canonical. List yang dihasilkan daripada blind crawl boleh mengekalkan parameter duplicate, redirect atau staging route.
Google cuba crawl exact absolute URL yang diberikan. Selaraskan protocol, hostname, path casing, trailing-slash policy dan locale path dengan final Canonical.
- Absolute fully qualified production HTTPS URL.
- Preferred Canonical version dengan hostname, path dan locale yang betul.
- Expected HTTP 200 dengan meaningful visible content.
- Dibenarkan robots.txt dan tidak protected by login.
- Tiada robots meta atau X-Robots-Tag noindex.
- Bukan duplicate parameter, sort, filter, session atau tracking variation.
- Dicapai melalui sekurang-kurangnya satu useful crawlable internal link.
- Sesuai sebagai landing page daripada search result.
- Disenaraikan sekali dalam correct child Sitemap.
- Dibuang automatik apabila publishing atau index decision berubah.
Exclude URL yang menyampaikan index story berbeza
| URL state | Sitemap treatment | Correct technical treatment |
|---|---|---|
| 301/308 permanent redirect | Exclude source; list final preferred URL sahaja | Old URL direct redirect ke closest equivalent |
| 302/307 temporary redirect | Biasanya exclude source daripada preferred inventory | Confirm URL mana perlu kekal indexed |
| 404/410 retired URL | Exclude | Kekalkan real removal response dan repair internal link |
| 5xx atau persistent timeout | Jangan tampilkan sebagai healthy inventory | Fix server atau application cause |
| noindex page | Exclude | Kekalkan crawlable sehingga Google process directive |
| Canonical duplicate | Exclude duplicate; include preferred representative | Align internal link, hreflang dan Canonical |
| Internal search atau weak filter | Exclude | Deliberate index/crawl rules ikut value dan scale |
| Staging, preview, admin atau account URL | Exclude | Protect private environment dan session dengan access control |
Patuhi protocol limit dan file scope
Satu Sitemap terhad kepada 50,000 URL atau 50 MB uncompressed. Sitemap index boleh reference sehingga 50,000 Sitemap file dan tertakluk kepada uncompressed size limit sama. Compression kurangkan transfer size, bukan protocol limit.
Gunakan UTF-8, entity-escape XML value dan stable public Sitemap URL. Melainkan disubmit melalui Search Console, Sitemap biasanya hanya meliputi descendant parent directory; root placement mengelakkan accidental scope restriction.
| Protocol item | Requirement | Audit check |
|---|---|---|
| URL set | Maksimum 50,000 URL dan 50 MB uncompressed | Count entry dan uncompressed byte sebelum release |
| Sitemap index | Maksimum 50,000 child Sitemap dan 50 MB | Child location resolve dan mengandungi intended host inventory |
| Encoding | UTF-8 dengan valid XML dan escaped entity | Parser menerima ampersand, non-Latin path dan namespace |
| Location | Stable accessible URL dalam valid scope | Root placement atau verified cross-site submission disengajakan |
| Page URL | Fully qualified absolute URL | Tiada relative path, development host atau mixed protocol |
| Ordering | Tiada ranking meaning | Jangan sort mengikut supposed importance |
Partition child Sitemap untuk diagnosis
Membahagikan small inventory kepada ratusan file menambah maintenance noise. Partition apabila kumpulan mempunyai template, owner, update pattern atau risk berbeza supaya Search Console filter menunjukkan actionable cohort.
SEOWithJack kini menggunakan satu root Sitemap index yang reference post, page dan category child Sitemap. Post Sitemap merangkumi English, Malay dan Simplified Chinese article URLs, manakala robots.txt menunjuk kepada satu root index. Pattern ini mudah diaudit dan masih boleh isolate content type.
| Possible child Sitemap | Useful apabila | Poor reason |
|---|---|---|
| Posts atau articles | Editorial URL berkongsi publishing dan refresh rule | Setiap bulan mesti ada permanent file |
| Pages atau services | Commercial/static template perlukan separate monitoring | URL count lebih kecil |
| Products atau categories | Template dan Canonical risk perlu cohort | Untuk claim higher priority |
| Locales | Team atau platform maintain language berasingan | Untuk elak proper hreflang relationship |
| Images, videos atau news | Specialized extension dan validation terpakai | Website mempunyai satu ordinary image |
| Legacy migration cohort | Temporary monitoring untuk moved URL didokumenkan | Redirect old URL dianggap preferred page |
Gunakan lastmod sebagai factual change record
Google mungkin menggunakan <lastmod> untuk crawl scheduling jika nilainya consistently dan verifiably accurate. Ia perlu mewakili last significant page change, bukan Sitemap generation time, deployment timestamp atau copyright year.
Google mengabaikan <priority> dan <changefreq>. Unauthenticated Sitemap Ping endpoint juga sudah deprecated dan memberi 404; gunakan stable Sitemap URL bersama Search Console, robots.txt atau Search Console API.
| Perubahan | Update lastmod? | Sebab |
|---|---|---|
| Rewrite main guidance dengan new evidence | Ya | Page berubah secara meaningful |
| Tambah atau correct important structured data | Ya | Google menganggapnya potentially significant |
| Ubah important internal link atau navigation context | Biasanya ya jika material | Discovery dan page relationship berubah |
| Fix satu typo atau compress same image | Biasanya tidak | Purpose dan information tidak berubah materially |
| Update footer, copyright atau global CSS sahaja | Tidak untuk semua content URL | Shared cosmetic build bukan content refresh |
| Republish unchanged text dengan tarikh hari ini | Tidak | Signal tidak accurate dan freshness misleading |
| Aggregator page berubah automatik | Hanya jika system boleh kira meaningful update dengan tepat | Omit lastmod jika confidence rendah |
Selaraskan Sitemap, status, Canonical dan index rule
| Signal | Preferred indexable URL | Retired URL | Duplicate URL | Private URL |
|---|---|---|---|---|
| HTTP response | 200 | 404/410 atau relevant 301/308 | 200 atau redirect ikut product need | 401/403 atau authenticated access |
| robots.txt | Allow crawl | Biasanya tiada rule khas | Allow jika Google perlu baca Canonical/noindex | Bukan security control |
| robots meta/header | Index dibenarkan | Tidak perlu untuk true removal | noindex hanya jika exclusion intended | Bukan security control |
| Canonical | Self-referential preferred URL | Tiada pada removed response | Ke equivalent representative jika sesuai | Tidak digunakan |
| Internal link | Terus ke preferred URL | Remove atau update | Prefer representative URL | Dalam authorized experience sahaja |
| Sitemap | Include | Exclude | Exclude nonpreferred duplicate | Exclude |
Multilingual Sitemap dan hreflang
XML ialah satu daripada tiga Google-supported hreflang method yang equivalent. Jika dipilih, setiap localized URL mempunyai <url> entry sendiri, dan setiap entry menyenaraikan dirinya serta semua alternate melalui identical xhtml:link annotation.
Jangan guna hreflang untuk membaiki translated navigation di sekeliling untranslated main content. Setiap locale URL perlu real indexable localized page dengan self-Canonical dan supported language code. Rujuk panduan multilingual SEO.
- Setiap indexable language URL mempunyai Sitemap URL entry.
- Setiap entry list dirinya dan semua alternate.
- Alternate set reciprocal dan identical.
- Semua href ialah absolute production URL.
- Setiap alternate HTTP 200 dan self-Canonical dalam locale itu.
- Language code menggunakan supported ISO language dan optional region.
- x-default hanya untuk deliberate fallback destination.
- Redirect, noindex, blocked atau untranslated variant tidak declared valid.
- HTML atau HTTP-header hreflang tidak contradict mapping.
Gunakan specialized extension hanya bila membantu discovery
| Extension | Useful untuk | Release gate |
|---|---|---|
| Image | Important image yang sukar ditemui, termasuk sesetengah JavaScript-reached asset | Image URL crawlable dan current supported tag sahaja |
| Video | Page dengan video sebagai main content | Landing page, thumbnail, player/content URL dan required field accessible |
| News | Eligible news publisher dengan current article inventory | Ikut current Google News Sitemap requirement |
| xhtml hreflang | Localized variant apabila Sitemap ialah chosen method | Setiap version list complete reciprocal set |
| Combined extensions | Page benar-benar layak untuk beberapa type | Declare namespace sekali dan validate XML |
| None | Ordinary page sudah discoverable dalam HTML | Jangan tambah unsupported metadata untuk appearance |
Submit location, bukan file contents
Search Console submission memberitahu Google lokasi hosted Sitemap; ia bukan upload XML kepada Google. File mesti kekal public dan fetchable. Successful submission bermaksud file boleh diprocess, bukan semua URL telah dicrawl, diindex atau rank.
| Method | Use case | Limit |
|---|---|---|
| Search Console Sitemaps report | Manual submission dengan fetch/parsing feedback | Perlu owner permission dan property scope tepat |
| robots.txt Sitemap line | Persistent public declaration | Gunakan fully qualified URL; processing tidak dijamin |
| Search Console API | Programmatic management untuk verified properties | Automation masih perlukan validation dan ownership |
| RSS/Atom dengan WebSub | Broadcast recent publishing change | Recent feed bukan full inventory |
| Deprecated Ping endpoint | Jangan guna | Google return 404 dan tiada useful signal |
Baca Search Console pada level betul
Gunakan Sitemaps report untuk fetch history dan parsing error. Gunakan Page indexing report yang difilter kepada “All submitted pages” atau specific Sitemap untuk melihat cohort. Gunakan URL Inspection bagi representative URL, bukan extrapolate whole site daripada satu page.
Gap antara submitted dan indexed URL bukan automatic error. Redirect, duplicate, noindex dan removal sepatutnya tiada dalam clean Sitemap; valid page masih mungkin perlukan diagnosis quality, Canonical selection atau processing.
| Soalan | Best evidence | Jangan simpulkan |
|---|---|---|
| Bolehkah Google fetch dan parse file? | Sitemaps report status, last read dan error | Success bermaksud semua URL indexed |
| Submitted cohort mana tidak indexed? | Page indexing filter by Sitemap | Setiap non-indexed URL ialah defect |
| Apa Google tahu tentang satu URL? | URL Inspection indexed data dan live test | Live test bermaksud sudah masuk index |
| Adakah generated output technically clean? | XML parse dan URL inventory crawl | Valid XML bermaksud URL decisions betul |
| Adakah discovery bertambah baik? | Comparable crawl/index cohort dan server log | Submission timestamp menyebabkan ranking gain |
robots.txt mesti berada di exact host root
File mesti tersedia sebagai lowercase /robots.txt pada top level. Rules hanya terpakai kepada protocol, hostname dan port sama. File pada https://example.com tidak mengawal HTTP, www, subdomain lain atau nonstandard port.
Translated subfolder berkongsi host robots file; translated subdomain perlukan file sendiri. CDN atau application perlu return plain UTF-8 text, bukan login page, HTML error template atau redirect chain.
| robots.txt location | Mengawal | Tidak mengawal |
|---|---|---|
| https://example.com/robots.txt | HTTPS URL pada example.com standard port | HTTP, www, subdomain atau custom port |
| https://www.example.com/robots.txt | www HTTPS host sahaja | Apex atau shop.example.com |
| https://shop.example.com/robots.txt | shop HTTPS subdomain sahaja | Main site atau subdomain lain |
| https://example.com/folder/robots.txt | Bukan root robots policy | URL di bawah /folder/ |
| https://example.com:8443/robots.txt | Host, protocol dan port itu sahaja | Standard HTTPS port |
Fahami group, matching dan supported field
| Element | Google behavior | Audit risk |
|---|---|---|
| User-agent | Mulakan group; Google pilih most specific matching token dan combine duplicate matching group | Generic group dianggap override specific group |
| Disallow | Block crawl request apabila path match | Broad prefix terkena useful URL |
| Allow | Permit exception dalam broader Disallow | Exception lebih pendek daripada blocking rule |
| Longest match | Most specific path menang; Allow menang same-length tie | Reviewer baca top-to-bottom sahaja |
| Path casing | Rule case-sensitive | Live URL capitalization berbeza |
| Wildcards | Google support * dan end-anchor $ | Regex assumption menghasilkan wrong match |
| Sitemap | Absolute location; tidak tied kepada user-agent group | Relative URL atau wrong host |
| Unsupported line | Google ignore | crawl-delay atau plugin directive dianggap universal |
Mulakan dengan robots.txt paling kecil
Jika semua public content boleh dicrawl, absent file atau empty applicable Disallow sudah bermaksud allow. Allow: / biasanya tidak perlu. Tambah rule hanya untuk pattern yang deliberate, kemudian test representative allowed dan blocked URL.
Simple service site boleh gunakan satu User-agent: * group, block internal search-result path dan list satu absolute Sitemap index. Jangan copy obsolete WordPress admin rule selepas route sudah tidak wujud atau block CSS, JavaScript dan image untuk public page.
- Setiap Disallow mempunyai documented crawl-management purpose.
- Rule match exact production case, slash dan parameter pattern.
- Specific Allow exception berfungsi mengikut longest-match.
- Tiada public page atau rendering asset blocked secara tidak sengaja.
- Tiada rule menjadi security, Canonical atau noindex substitute.
- Specific crawler group ikut published documentation crawler itu.
- Sitemap line menggunakan final absolute production URL.
- Staging rule tidak boleh disalin silently ke production.
Fahami kesan robots.txt response failure
robots.txt ialah operational infrastructure. Google biasanya cache sehingga 24 jam dan mungkin lebih lama jika refresh gagal. Perubahan memerlukan propagation time, manakala outage boleh mempengaruhi seluruh host.
| Response | Google-documented treatment | Operational response |
|---|---|---|
| 2xx | Process valid rules diterima | Check content type, encoding dan intended rules |
| 3xx | Follow sekurang-kurangnya lima redirect; selepas itu seperti 404 | Serve direct di canonical host root jika boleh |
| 4xx kecuali 429 | Anggap tiada robots file dan tiada crawl restriction | Jangan guna 401/403 untuk throttle |
| 429 | Dikendalikan berbeza daripada ordinary 4xx | Review capacity dan rate limiting |
| 5xx | Pada awalnya stop crawling 12 jam; mungkin guna last good version sambil retry | Urgent host reliability issue |
| DNS/network failure | Dianggap server error | Fix DNS, TLS, edge atau availability |
| Lebih 500 KiB | Content selepas limit diabaikan | Consolidate rules dan restructure pattern |
| Cached old version | Boleh kekal sekitar 24 jam atau lebih semasa failure | Allow propagation dan verify delivery |
robots.txt tidak boleh remove atau secure content
| Keperluan | Correct method | Kenapa robots.txt gagal |
|---|---|---|
| Remove accessible HTML page dari Google | crawlable meta robots noindex | Blocked crawler tidak boleh baca page rule |
| Remove accessible PDF/non-HTML | crawlable X-Robots-Tag noindex header | HTML meta tiada dan block sembunyikan header |
| Urgently hide result sementara durable fix disiapkan | Search Console Removals bersama durable noindex/removal/auth | Robots change sahaja tidak immediate atau durable |
| Protect customer, staging atau confidential data | Authentication, authorization dan network control | robots.txt public, advisory dan expose named path |
| Retire content tanpa replacement | 404/410 dan remove dari link/Sitemap | Disallow biarkan URL unresolved |
| Move content permanently | 301/308 ke closest equivalent | Disallow ganggu processing move |
| Consolidate duplicate | Canonical signal atau permanent redirect | Block menghalang Canonical signal dibaca |
Generate mengikut WordPress, static site dan application
| Platform | Source of truth | Release risk |
|---|---|---|
| WordPress core | Published public content dan native Sitemap jika sesuai | SEO/cache/multilingual plugin create second conflicting index |
| WordPress dengan SEO plugin | Satu selected provider ikut final Canonical settings | Submit native dan plugin inventory tanpa compare |
| Static site generator | Build manifest untuk public indexable route | Setiap build tukar lastmod atau include utility page |
| Headless CMS | Published record digabung route, locale dan Canonical state | Unpublished atau orphan record bocor ke XML |
| Ecommerce/application | Product/category availability dan index-decision service | Facet, session, sort parameter dan soft-deleted record |
| Multiple hosts | Per-host inventory dengan verified cross-submission sahaja | Satu robots file dianggap control semua subdomain |
Protect WordPress-to-static migration
- Export old Sitemap inventory, public crawl, Canonical, hreflang dan index directive.
- Map setiap old URL kepada unchanged route, true equivalent redirect atau justified 404/410.
- Protect staging dengan authentication dan exclude hostname daripada production Sitemap.
- Generate Sitemap daripada final published route manifest, bukan blind crawl.
- Compare old dan new preferred inventory ikut content type dan locale.
- Remove redirect, noindex, error, search result dan non-Canonical variant.
- Validate XML, child index, limit, response dan production hostname.
- Publish redirect, page, Sitemap dan robots.txt dalam release sama.
- Confirm production robots tidak mewarisi site-wide staging Disallow.
- Crawl semua old URL dan new Sitemap URL selepas launch.
- Submit stable Sitemap index dalam production Search Console property.
- Monitor submitted cohort, selected Canonical, server error dan organic landing page.
Diagnose symptom mengikut urutan
| Symptom | Likely layer | Evidence pertama |
|---|---|---|
| Sitemap “Could not fetch” | Access, DNS, redirect, content type atau property scope | Direct HTTP response dan report detail |
| XML parsing error | Encoding, escaping, namespace atau malformed markup | Raw response dan XML validator |
| Submitted URL blocked robots.txt | Inventory dan crawl policy conflict | Exact entry dan matching rule |
| Submitted URL noindex | Publishing/index decision conflict | Rendered/head response dan generator logic |
| Submitted URL redirect | Old atau wrong Canonical inventory | Redirect destination dan route manifest |
| Different Google-selected Canonical | Duplicate atau inconsistent signal | Page pair, links, Canonical, hreflang dan Sitemap |
| Important URL tiada Sitemap | Generator, publishing state atau partition error | Source record dan child assignment |
| Large unsubmitted URL set | Facet, parameter, old route atau crawler trap | Page indexing filter, crawl dan log |
| robots changed tetapi test nampak old | Caching atau edge delivery | Response header, host variant dan elapsed time |
| Traffic berubah selepas submit | Bukan proof Sitemap causation | Query, page, release, demand dan SERP cohort |
Gunakan release gate, bukan visual spot-check
- robots.txt dan semua Sitemap HTTP 200 pada production host.
- File UTF-8 dan XML parse dengan setiap namespace.
- Child location absolute, unique dan fetchable.
- Setiap listed page ialah 200, indexable, Canonical dan internally linked.
- Tiada redirect, error, login, staging atau non-Canonical URL.
- Locale cluster complete dan reciprocal apabila Sitemap hreflang digunakan.
- lastmod omitted atau tied kepada meaningful source-controlled change.
- robots rules pass allowed, blocked dan exception test.
- Required rendering assets crawlable.
- Sitemap index declared sekali dengan exact absolute URL.
- Old WordPress URL ikut approved redirect map.
- Search Console boleh fetch dan cohort monitoring mempunyai owner.
Mitos Sitemap dan robots.txt
- “Semua Sitemap URL akan indexed.” Ia hint, bukan guarantee.
- “Sitemap ganti internal link.” Orphan page masih architecture problem.
- “priority 1.0 naikkan ranking.” Google ignore priority.
- “changefreq=daily paksa daily crawl.” Google ignore changefreq.
- “Fresh lastmod guarantee recrawl hari ini.” Ia berguna hanya jika accurate.
- “Old Ping URL percepat submission.” Google deprecated dan return 404.
- “robots.txt remove page dari Cari.” Blocked URL masih boleh indexed tanpa fetched content.
- “Disallow protect confidential path.” File public dan tiada authorization.
- “Block noindex page untuk extra certainty.” Block boleh halang noindex dibaca.
- “Satu robots file control www, subdomain dan HTTP.” Scope ikut protocol, host dan port.
- “Lebih banyak Disallow sentiasa save crawl budget.” Complexity boleh create bigger failure.
- “Successful submission bermaksud SEO complete.” Ia file processing, bukan page quality atau ranking.
Soalan lazim
Perlukah setiap website XML Sitemap?
Tidak. Small comprehensively linked site mungkin ditemukan tanpa Sitemap. Namun ia berguna untuk launch, migration, multilingual inventory, specialized media dan monitoring.
Adakah Sitemap guarantee crawl atau index?
Tidak. Ia hint. URL masih perlu accessible, indexable, canonical, useful dan dipilih oleh Google.
Patut redirect atau noindex URL berada dalam Sitemap?
Tidak. Preferred inventory perlu exclude redirect, removal, noindex dan non-Canonical duplicate.
Berapa URL dalam satu Sitemap?
Maksimum 50,000 URL atau 50 MB uncompressed, mana yang dicapai dahulu. Gunakan child Sitemap dan index jika lebih besar.
Perlukah lastmod pada setiap URL?
Hanya jika system mampu memberi accurate last significant change date. Omit apabila confidence rendah.
Adakah priority dan changefreq membantu?
Tidak. Google menyatakan kedua-duanya diabaikan.
Patut noindex URL juga blocked?
Tidak. Google perlu crawl untuk membaca meta robots atau X-Robots-Tag noindex.
Bolehkah robots.txt protect staging?
Tidak. Gunakan authentication atau network access. Public Disallow bersifat advisory.
Di mana host robots.txt dan Sitemap?
robots.txt mesti di exact host root. Root-level Sitemap memberi scope mudah dan boleh declared dalam robots.txt serta Search Console.
Bagaimana monitor Sitemap?
Semak fetch/parsing dalam Sitemaps report, filter Page indexing ikut Sitemap, inspect representative URL dan compare cohort.
Rujukan rasmi
- Google: build and submit a sitemap
- Google: sitemaps overview
- Sitemaps.org protocol
- Google Search Console: Sitemaps report
- Google Search Console: Page indexing report
- Google: introduction to robots.txt
- Google: create and submit a robots.txt file
- Google: how Google interprets robots.txt
- IETF RFC 9309: Robots Exclusion Protocol
- Google: block indexing with noindex
- Google robots meta specifications
- Google: localized page versions and hreflang
- Google: combine Sitemap extensions
- Google: image Sitemaps
- Google: Sitemap ping endpoint deprecation and lastmod
Kongsikan live Sitemap index, robots.txt, Search Console state dan URL pattern. Jack boleh membandingkan inventory dengan Canonical, response, locale route dan migration redirect.



Evolusi Ranking Google: Daripada PageRank ke Modern Cari Systems2 September 2026
Kandungan AI dan SEO Google: Apa Dibenarkan, Apa Dianggap Spam dan Cara Publish2 September 2026
Google Florida Update 2003: Fakta, Teori dan Lesson SEO2 September 2026