Organic visibility bergantung pada beberapa sistem berasingan. Google perlu discover URL, memilih untuk crawl, menerima response yang berguna, render JavaScript jika perlu, process page, memilih Canonical dan akhirnya mempertimbangkan page relevan untuk sesuatu query. Lulus satu stage tidak menjamin stage seterusnya.
Jangan bermula dengan “request indexing.” Kenal pasti sama ada masalah berlaku pada discovery, crawler access, server response, rendering, indexability, duplication, Canonical selection, content value atau query relevance. Setiap stage memerlukan evidence dan fix berbeza.
Cari pipeline daripada URL ke result
Google menerangkan tiga stage besar—crawling, indexing dan serving—tetapi diagnosis teknikal perlu memisahkan transition di dalamnya. URL boleh diketahui tanpa difetch, difetch tanpa rendered dengan betul, diproses tanpa dipilih sebagai Canonical, atau diindeks tanpa dipaparkan untuk query yang diperiksa.
| Stage | Aktiviti | Evidence | Kesimpulan salah |
|---|---|---|---|
| Discovery | Google mengetahui URL wujud | Internal link, Sitemap, redirect, external link | “Dalam Sitemap bermaksud sudah crawl” |
| Crawl scheduling | Google memilih bila hendak request | Logs, Crawl Stats, last crawl | “Known bermaksud crawl segera” |
| Fetching | Crawler request URL dan resources | HTTP response, headers, timing | “Browser okay bermaksud Googlebot okay” |
| Rendering | HTML/JavaScript menghasilkan document | Source dan rendered HTML | “Google melihat semua client state” |
| Index processing | Content, directives dan duplicate dinilai | Page Indexing, URL Inspection | “200 menjamin index” |
| Canonical selection | Satu representative dipilih | Declared dan selected Canonical | “Canonical ialah directive” |
| Serving | Indexed page mungkin dipilih untuk query | Performance by page/query | “Indexed bermaksud rank” |
Mulakan dengan technical eligibility
Keperluan minimum Google ialah Googlebot tidak disekat, page berfungsi dengan HTTP success dan mempunyai indexable content. Ini hanya menjadikan indexing mungkin; ia tidak menjamin crawl, index atau ranking.
- Exact preferred URL public dan boleh resolve.
- Googlebot dibenarkan oleh robots.txt untuk host itu.
- Final page memberi stable HTTP 200.
- Primary content lengkap dan berguna pada mobile.
- Tiada unintended noindex dalam meta atau X-Robots-Tag.
- Essential content tidak memerlukan login, consent atau interaction.
- Declared Canonical valid dan selari.
- Content tidak melanggar spam atau legal policies.
Discovery memerlukan crawlable path yang kekal
Google menemui URL melalui known pages, standard links, redirects, Sitemaps dan external links. Sitemap ialah inventory, bukan pengganti navigation. Important page perlukan normal anchor link daripada relevant hub tanpa form submission atau script event.
Gunakan panduan site architecture dan panduan internal linking untuk menghubungkan services, topic hubs dan supporting articles.
| Discovery source | Peranan | Audit question |
|---|---|---|
| Normal anchor dengan destination | Primary repeatable path | Boleh dicapai daripada indexable hub? |
| XML Sitemap | Preferred URL inventory hint | Final, Canonical dan 200? |
| Redirect | Old path menunjuk kepadanya | Relevant dan direct? |
| External link | Independent discovery/reference | Resolve tanpa access error? |
| JS-inserted anchor | Boleh selepas rendering | Real href dalam rendered HTML? |
| Button/onclick/fragment route | Unreliable | Boleh jadi standard route dan anchor? |
Discovered tidak bermaksud dicrawl segera
Selepas discovery, Google memilih secara algorithmic URL yang hendak dicrawl, frequency dan volume yang host boleh tanggung. Scheduling dipengaruhi crawl demand, host capacity, change, importance, duplication dan URL inventory. Tiada fixed frequency untuk semua pages.
“Discovered – currently not indexed” bermaksud URL diketahui tetapi belum difetch. Kuatkan meaningful internal importance, server reliability dan inventory quality sebelum repeated manual submission.
| Condition | Possible effect | Response |
|---|---|---|
| Relevant internal links | Clearer importance/discovery | Link daripada appropriate hubs |
| Accurate Sitemap/lastmod | Cleaner scheduling hint | Update selepas substantive change |
| Stable fast server | Higher safe capacity | Monitor latency dan 5xx |
| Banyak duplicate/filter URL | Activity tersebar | Control generation dan Canonicals |
| Low demand/unchanged content | Less frequent recrawl | Jangan cipta meaningless update |
| Site move/launch | Temporary demand change | Direct redirects dan new Sitemap |
robots.txt mengawal request, bukan indexing
robots.txt memberitahu compliant crawlers URL yang boleh diminta pada protocol, host dan port tepat. Ia bukan index-removal atau security tool. Blocked URL masih boleh diketahui dan muncul dengan limited information kerana Google tidak boleh crawl untuk membaca content atau noindex.
Untuk keluarkan public page daripada Cari, benarkan crawling dan gunakan noindex, pulangkan 4xx yang betul, atau lindungi private content dengan authentication. Lihat panduan Sitemap dan robots.txt.
| Goal | Control | Sebab |
|---|---|---|
| Kurangkan crawling URL space | robots.txt + generation controls | Hentikan permitted crawler requests |
| Exclude accessible HTML | Robots meta noindex | Crawler boleh membaca rule |
| Exclude PDF/non-HTML | X-Robots-Tag noindex | Rule dalam response header |
| Remove deleted URL | 404/410 | Content tidak wujud |
| Protect confidential content | Authentication/authorization | Cegah unauthorized retrieval |
HTTP response menentukan processing
| Response | Maksud | Audit action |
|---|---|---|
| 200 | Content boleh diproses; index tidak dijamin | Semak useful content dan directives |
| 301/308 | Permanent move signal | Satu relevant final 200 |
| 302/303/307 | Temporary routing | Source patut kekal long-term |
| 304 | Reuse previous representation | Validators ikut real changes |
| 404/410 | Resource tidak wujud | Kekalkan jika intentional |
| 429 | Server overload signal | Control load dan retries |
| 5xx/network/DNS | Host tidak reliable | Urgent jika sustained |
| 200 dengan empty/error content | Boleh jadi soft 404 | Honest status atau restore content |
Browser berjaya belum mencukupi
CDN, firewall, bot protection, geolocation, cookies, device detection, redirects dan intermittent origin failures boleh memberi Googlebot Smartphone response berbeza daripada user biasa.
- Exact requested URL, final URL dan semua redirect hops.
- HTTP status, headers, content type dan response time.
- robots.txt result untuk host dan crawler betul.
- Source HTML sebelum JavaScript.
- Rendered HTML dan essential resources.
- Mobile content, metadata, Canonical dan schema parity.
- Repeated samples untuk expose intermittent failure.
- Server logs mengesahkan crawler sampai ke application.
Rendering ialah diagnostic layer berasingan
Google menggunakan recent Chromium dan boleh execute JavaScript, tetapi rendering menambah dependency kepada scripts, APIs, CORS, CSP, client routing, hydration dan resource availability. Essential content jangan menunggu click, swipe, typing, nonessential consent atau scroll event.
Source HTML yang sudah mengandungi primary content, headings, crawlable links dan stable metadata lebih resilient. SSR atau static generation membantu hanya jika delivered HTML betul dan hydration tidak menggantikannya dengan error state.
| Layer | Banding | Failure |
|---|---|---|
| HTTP response | Status, headers, raw body | 200 app shell tanpa content |
| Source HTML | Title, H1, links, Canonical, robots | Metadata selepas failed API |
| Rendered DOM | Final content dan anchors | Hydration buang text |
| Resources/API | JS, CSS, images, data | Blocked API beri empty page |
| Indexed view | Last processed version | Live fix belum diproses |
Mobile content ialah indexing baseline
Google menggunakan mobile version untuk indexing dan ranking. Responsive design biasanya paling mudah, tetapi semua configuration perlu equivalent primary content, metadata, schema, images dan index controls pada mobile.
Mobile accordion boleh digunakan jika content wujud dalam rendered page. Primary content yang hanya muncul selepas user interaction tidak reliable.
- Primary copy, headings dan important links equivalent.
- Title, description, robots dan Canonical selari.
- Structured data menerangkan visible entities sama.
- Images ada useful alt dan accessible URLs.
- Tiada mobile-only noindex atau blocking rule.
- Primary lazy content tidak perlu interaction.
Indexing ialah analysis dan selection
Selepas crawl dan render, Google process text, images, video, title, alt, structured data, language dan signals. Ia menilai primary content, detect duplicates dan mungkin menyimpan selected Canonical bersama cluster. Tidak semua processed page akan diindeks.
Valid 200 page boleh kekal unindexed kerana duplicate, incompatible Canonical, soft 404, nilai tidak distinct atau tidak dipilih sistem Google. Repeated request tidak menjadikan page lebih berguna.
| Gate | Healthy | Failure |
|---|---|---|
| Index permission | Tiada unintended noindex | Meta/header conflict |
| Primary content | Useful dan template-specific | Empty, thin, repeated |
| Canonical | Final 200 preference | Redirect/noindex target |
| Language/locale | Page dan cluster agree | Partial translation |
| Mobile/render parity | Essential content survives | API/interaction hides content |
| Site context | Relevant hubs dan links | Orphan/near duplicate |
Canonical selection berlaku dalam indexing
Google cluster similar pages dan memilih representative. Redirect serta rel=canonical ialah strong signals; Sitemap lebih lemah. Internal links, HTTPS dan content similarity turut mempengaruhi. Google boleh memilih Canonical lain jika target tidak equivalent atau evidence conflict.
Rujuk panduan Canonical dan redirects sebelum menukar URL atau consolidate pages.
| State | Maksud | Action |
|---|---|---|
| Alternate with proper Canonical | Expected consolidation | Sahkan deliberate |
| Duplicate without selected Canonical | Variants dikelompok tanpa clear declaration | Align preference signals |
| Google chose different Canonical | URL lain lebih representative | Compare content dan signals |
| Page with redirect | Source bukan destination indexable | Inspect final target |
| Canonical target not indexed | Target ada issue sendiri | Diagnose first failed stage |
Indexing dan serving ialah outcome berbeza
Indexed page layak muncul, bukan dijamin rank. Apabila user search, Google menilai relevance, quality, language, location dan device. Cari features juga berubah mengikut query.
Jika URL indexed tetapi tiada impressions, audit search intent, demand, competition, usefulness, internal context dan measurement—bukan crawling secara default. Gunakan panduan search intent dan panduan measurement.
| Evidence | Membuktikan | Tidak membuktikan |
|---|---|---|
| Indexed | Google menyimpan selected representation | Rank target query |
| Impression | Result dipaparkan | Click/conversion |
| Click | User memilih result | Visit useful |
| Organic session | Analytics rekod visit | Exact GSC match |
| Lead/conversion | Business action berlaku | SEO sahaja menyebabkannya |
Fahami Page Indexing tanpa mengejar 100 peratus
| State | Maksud | Priority |
|---|---|---|
| Discovered – not indexed | Known tetapi belum crawl | Important cohorts, discovery dan inventory |
| Crawled – not indexed | Fetched tetapi tidak selected | Value, duplicate, Canonical, soft 404 |
| Excluded by noindex | Rule dibaca | Fix hanya jika patut public |
| Blocked by robots | Content tidak boleh fetch | Fix jika perlu access/noindex |
| Page with redirect | Source ke tempat lain | Expected jika intentional |
| Alternate/duplicate | Canonical lain selected | Expected untuk variants |
| Server error | 5xx | Urgent jika widespread |
| Soft 404 | Content error-like walaupun 200 | Restore atau honest status |
Gunakan Search Console mengikut tugas sebenar
| Tool | Best use | Limit |
|---|---|---|
| Page Indexing | Patterns/totals known URLs | Examples ialah subset |
| URL Inspection index data | Last processed state satu URL | Berdasarkan last cycle |
| Live test | Current fetch/render selepas fix | Tiada duplicate clustering/guarantee |
| Sitemaps | Fetch/parsing dan submitted inventory | Hint, bukan approval |
| Crawl Stats | Host, requests, responses, type/purpose | Advanced; examples tidak complete |
| Performance | Queries, pages, countries, devices | Canonical/privacy affect totals |
| Removals | Temporary hide owned URL | Bukan permanent solution |
Baca URL Inspection dalam urutan betul
- Inspect exact preferred URL termasuk protocol, host dan path.
- Semak sama ada URL diketahui dan catat last crawl.
- Sahkan crawl allowed, fetch dan response.
- Sahkan indexing allowed dan robots rules.
- Banding user-declared dan Google-selected Canonical.
- Review discovery dan Sitemap association.
- Run live test untuk current response/render.
- Jangan keliru indexed evidence dengan live result.
- Baiki generating template/routing rule.
- Request indexing sekali selepas meaningful fix dan monitor cohort.
Crawl Stats dan logs menjawab soalan berbeza
Crawl Stats merumuskan host, response, file type, purpose dan Googlebot type; example URLs tidak complete. Server/CDN logs memberi request-level evidence tetapi tidak membuktikan index atau rank.
| Soalan | Evidence | Caution |
|---|---|---|
| Googlebot sampai ke host? | Crawl Stats + verified logs | Verify genuine Googlebot |
| Response mana meningkat? | Response groups + logs | Spike mungkin launch |
| Pattern mana consume requests? | Full logs | Samples bukan inventory |
| Redirect tambah requests? | Hop logs/crawler | Setiap hop dikira |
| URL indexed? | URL Inspection index data | Crawl log bukan index |
| Visibility berubah? | Performance by cohort | GSC dan analytics berbeza |
Kebanyakan site tidak perlukan advanced crawl-budget project
Google meletakkan crawl-budget management untuk site sangat besar atau frequently updated. Search Console menyatakan site bawah kira-kira seribu pages biasanya tidak perlu Crawl Stats-level optimization. Untuk service site biasa, technical clarity dan content value lebih penting.
Large ecommerce, marketplace, publisher dan faceted site mungkin perlukan deeper control. Crawl budget menggabungkan crawl capacity dan crawl demand.
| Condition | Priority | Action |
|---|---|---|
| Small service site | Low budget concern | Fix links, errors, orphans, duplicates |
| New site | Discovery/value | Strong hubs dan clean Sitemap |
| Large faceted ecommerce | Inventory/logs | Control combinations dan empty states |
| Publisher | Freshness/capacity | Accurate updates dan stable server |
| Migration | Temporary demand/redirects | One-hop map dan both hosts |
| Sustained 5xx/network | Capacity emergency | Fix infrastructure dahulu |
Kawal URL inventory pada source
Jangan cuba membersihkan infinite filters, session IDs, calendars, search pages atau duplicate paths melalui Search Console berulang. Cegah unnecessary URL daripada dijana atau linked, tetapkan Canonical, pulangkan honest status dan expose purposeful routes sahaja.
Current SEOWithJack build memisahkan 710 HTML documents daripada 459 indexable Sitemap URLs, preserve 27 original WordPress article routes dan audit semua generated internal targets. Public files, redirect responses dan indexable Canonicals ialah inventory berbeza.
| Inventory | Include | Exclude |
|---|---|---|
| Crawlable graph | Useful public pages/resources | Broken/event-only routes |
| Indexable Canonicals | Unique final Cari pages | Redirect/noindex/error/duplicate |
| XML Sitemap | Preferred 200 Canonicals | Parameters/retired URLs |
| Redirect map | Every old source/outcome | Unknown catch-all |
| QA crawl | Templates, locales, edge cases | Laman Utamapage-only sample |
Release gate untuk crawl dan index
- Setiap intended page ada stable final URL dan HTTP 200.
- robots.txt tersedia dan membenarkan required crawling.
- Staging noindex tidak masuk production.
- Source/rendered HTML ada equivalent content dan metadata.
- Mobile mempunyai content, links dan schema sama.
- Canonical, hreflang, links dan Sitemap guna final URLs.
- Removed routes beri relevant redirect, 404 atau 410.
- Redirect tanpa loop atau avoidable chain.
- 404, forms, phone dan WhatsApp berfungsi.
- Representative pages diuji selepas launch.
- Monitoring meliputi server, indexing cohorts dan conversions.
Workflow troubleshooting berdasarkan stage
| Symptom | First stage | Evidence |
|---|---|---|
| URL tiada dalam report | Discovery | Links/Sitemap |
| Discovered belum crawl | Scheduling/inventory | Importance, server, URL growth |
| Crawl failed | Fetching | Status, DNS, firewall, logs |
| Live page kosong | Rendering | Source vs rendered/API |
| Crawled not indexed | Indexing | Value, soft 404, noindex, duplicate |
| Canonical lain selected | Clustering | Three-URL comparison |
| Indexed tiada impressions | Serving/relevance | Page-query intent |
| Drop selepas migration | Routing/transfer | Cohorts, redirects, logs |
| Traffic ada, leads turun | Conversion | Forms/calls/events |
Mitos crawling dan indexing
- “Sitemap menjamin crawl dan index.”
- “Repeated request indexing lebih cepat.”
- “HTTP 200 bermaksud indexed.”
- “Indexed mesti rank untuk target keyword.”
- “robots.txt cegah URL muncul.”
- “Google sentiasa tunggu semua JavaScript.”
- “Live test tunjuk Canonical indexed.”
- “Semua non-indexed URL ialah error.”
- “Semua discovered URL mesti 100% indexed.”
- “Small business perlukan crawl-budget manipulation.”
- “Tukar publish date menjamin recrawl.”
- “Server logs membuktikan index dan rank.”
Soalan lazim
Apa beza crawling dan indexing?
Crawling ialah request URL/resources. Indexing ialah memproses content, directives, duplicates dan signals untuk menentukan apa yang disimpan.
Berapa lama Google index page?
Tiada fixed time. Recrawl boleh mengambil hari hingga minggu dan request tidak menjamin index.
Adakah Sitemap menjamin index?
Tidak. Ia membantu discovery; page masih perlu lulus access, processing, Canonical dan quality.
Mengapa crawled tetapi tidak indexed?
Duplicate, selected Canonical lain, thin value, soft 404, rendering atau selection issue.
Boleh indexed jika robots.txt block?
URL kadang-kadang boleh muncul berdasarkan external information kerana robots.txt block fetch, bukan discovery.
Adakah noindex menjimatkan crawl budget?
Noindex perlu dicrawl untuk dibaca dan boleh dicrawl semula. Ia index control, bukan fix infinite URL.
Perlu request indexing selepas setiap edit?
Tidak. Gunakan untuk sedikit important URL selepas meaningful fix; publishing biasa guna links dan Sitemap.
Index data vs live test?
Index data ialah last processed view; live test ialah current fetch/render dan tidak menilai semua indexing decisions.
Perlu optimize crawl budget?
Biasanya tidak untuk small/medium service site; perlu jika large/faceted inventory benar-benar melambatkan discovery.
Apa semak dahulu bila traffic jatuh?
Pisahkan availability, index coverage, Canonical, ranking/query demand dan conversion tracking.
Rujukan rasmi
- Google: how Cari works
- Google crawling and indexing overview
- Google Cari technical requirements
- Search Console Page indexing report
- Search Console URL Inspection
- Google: request a recrawl
- Google: JavaScript SEO basics
- Google: crawlable link best practices
- Google: robots.txt introduction
- Google: robots meta and X-Robots-Tag
- Google: canonicalization explained
- Google: build and submit a Sitemap
- Google: HTTP status codes for crawling
- Google: troubleshoot crawling errors
- Search Console Crawl Stats report
- Google: crawl budget management
- Google: mobile-first indexing best practices
- Search Console data limitations
Kongsikan affected URL cohort, Search Console state, recent release dan server evidence. Jack boleh memisahkan discovery, fetching, rendering, indexing, Canonical dan relevance sebelum mencadangkan perubahan.



Evolusi Ranking Google: Daripada PageRank ke Modern Cari Systems2 September 2026
Kandungan AI dan SEO Google: Apa Dibenarkan, Apa Dianggap Spam dan Cara Publish2 September 2026