Document type: Protocol | Version: 1.0 | Published: 28 July 2026
Executive Abstract
A protocol for distinguishing search crawlers, training crawlers, user-triggered fetchers, robots controls, log evidence, and AI referral attribution.
Definisi
Protocol ini mencegah pencampuran fungsi crawler. Search crawler, training crawler, user-triggered fetcher, dan advertising crawler dapat memiliki tujuan serta kontrol yang berbeda. Pengaturan robots.txt harus dipetakan terhadap tujuan bisnis dan bukti log server.
Ruang Lingkup
- OAI-SearchBot, GPTBot, ChatGPT-User, Googlebot, dan user agent relevan lainnya.
- Robots.txt, server log, CDN log, status response, crawl permission, dan referral tagging.
- Pemisahan discovery, training permission, live fetch, dan traffic attribution.
Batas dan Pengecualian
- Tidak menjamin inclusion hanya karena crawler diizinkan.
- Tidak menyimpulkan training dari satu hit user agent.
- Tidak menyamakan referral dengan citation atau recommendation.
Model Operasional
Access Policy
Dokumentasikan user agent, directive, path scope, dan tujuan izin atau blokir.
Server Verification
Gunakan log dan verifikasi IP sesuai dokumentasi provider bila tersedia.
Index and Surface Check
Pisahkan crawl success dari index eligibility dan answer inclusion.
Referral Attribution
Gunakan referrer, UTM, landing page, dan session behavior secara konsisten.
Change Control
Setiap perubahan robots harus memiliki alasan, tanggal, owner, dan rollback plan.
Langkah Implementasi atau Pengujian
- 1. Inventaris semua AI-related user agents yang terlihat di log.
- 2. Bandingkan robots.txt dengan kebijakan yang dimaksud.
- 3. Uji status response pada halaman representatif.
- 4. Pisahkan report crawl, inclusion, citation, referral, dan conversion.
- 5. Review setelah perubahan platform atau robots policy.
Evidence Requirements
- Snapshot robots.txt.
- Log request dengan timestamp dan user agent.
- HTTP response code dan canonical target.
- Analytics referral record.
- Dokumentasi resmi provider.
Governance dan Versioning
Perubahan akses crawler adalah keputusan governance, bukan sekadar setting teknis. Owner konten, legal, security, dan analytics perlu memahami konsekuensinya.
Keterbatasan
- User agent dapat dipalsukan.
- Tidak semua provider memberi IP verification yang sama.
- Absence dari log tidak membuktikan absence dari training atau indeks historis.
Sumber Primer
GEO.or.id operates as an independent research and methodology platform. This document does not provide a ranking guarantee, legal opinion, or permanent visibility claim.
Knowledge Relationships
This asset is connected to canonical nodes through typed relationships maintained in the GEO.or.id knowledge registry.
- Is Part Of: Protocols
- Related To: Agent Action Authorization Protocol
- Related To: Agent2Agent Protocol and Agent Cards
- Related To: Agentic Product Identity Protocol
- Related To: Graph
- Related To: System Map
