Study 01 · Global Digital Authority Benchmark Series · Australia 2026 · v1.5

How Australian businesses govern AI-crawler access

A structured benchmark of 409 publicly identifiable Australian business websites across ten industry groups, measuring AI-crawler access by the functional purpose of what is restricted — not by a single block/open count.

Research instrumentPTODA C01 Crawler v1.5.1 — deterministic robots.txt scanner
Scan date28 June 2026 (wired-confirmed)
Sample409 domains · 10 sectors · 317 with readable robots.txt policy
5.0%
whole-site exclusion of AI retrieval crawlers — the strict measure, of 317 policy-observed domains

A binary count records 42.6% of policy-observed Australian domains as “blocking AI.” Read by the functional purpose of the restricted path, whole-site exclusion is just 5.0% and restriction reaching primary public content is 11.7%.

Most restriction falls on operational and secondary paths — carts, search endpoints, APIs, admin routes — not on the primary content an AI system retrieves to answer a question.

Key findings

Restriction read at three thresholds

The same robots.txt policy data, read at three levels of stringency. Each measure includes the one above it and adds a less-substantive category. Strict and meaningful are final; the expanded measure is a conservative upper bound pending the unknown-path review.

5.0%
Strict restriction
16 of 317. Whole-site exclusion of AI crawlers (Disallow: /).
11.7%
Meaningful restriction
37 of 317. Restriction reaching primary public content, or whole-site.
51.1%
Expanded restriction
162 of 317. Any restriction including secondary and operational paths. Conservative upper bound.

The central finding: whole-site exclusion of AI retrieval crawlers is uncommon (5.0%), and meaningful restriction of public content is modest (11.7%). Most robots.txt activity classified as restrictive concerns functional, secondary, or infrastructure-layer paths rather than outright exclusion of AI crawlers from substantive content.

Five-class distribution

Access classified by functional purpose

Every policy-observed domain is classified by its highest substantive restriction, from fully open to whole-site blocked. Of 317 policy-observed Australian domains:

Five-class access distribution (of 317 policy-observed domains)
Fully open
44.2%
Functionally open
4.7%
Secondary restricted
39.4%
Primary restricted
6.6%
Whole-site blocked
5.0%

Just under half of policy-observed domains (48.9%) are fully or functionally open — they restrict, at most, operational infrastructure. A further 39.4% restrict only secondary content. Primary-content restriction (6.6%) and whole-site blocking (5.0%) together account for the 11.7% meaningful-restriction figure.

Recovery protocol & replication

Bounded recovery, independently replicated

Crawl-based measurement is subject to transient failure. The study used a bounded three-attempt recovery protocol with randomised order, and the Australian measurement was crawled twice on independent connections.

305
Initial observed
Attempt 1, wired connection.
+12
Recovered
Attempts 2–3, re-crawling unresolved domains only.
317
Policy-observed
Final denominator for all policy figures.

Independently replicated. Two crawls on separate network connections produced an identical policy denominator (317) with identical strict and meaningful restriction membership — the same 16 and same 37 domains. The wired run is canonical (c01_au_409_v15_wired); the mobile run is retained as replicated verification evidence. This establishes the result as reproducible rather than connection-dependent.

Infrastructure layer

Access outcomes — not crawler policy

Declared robots.txt policy is separated from infrastructure non-response. A 403, timeout or non-resolving response is an access outcome, not evidence of crawler policy. These 92 domains are reported here and excluded from every policy-layer figure.

38
Access denied (HTTP 403/401/429)
9.3% of the 409 domains approached. The edge refused the request; no crawler policy observable.
54
Unscannable
13.2% of domains. DNS (27), timeout (19), connection (4), client error (4). No readable response.
317
Policy observed
77.5% of domains returned a readable robots.txt. The denominator for all policy-layer findings.

Why this matters: a domain that denies the crawler at the infrastructure layer has not expressed an AI-crawler policy — it has prevented one from being read. Reporting these 92 domains separately keeps the policy-layer figures based only on directly observed robots.txt behaviour. Across the two independent crawls the excluded set was identical in membership.

Block origin

Intentional vs infrastructure-imposed

Of the 135 domains disallowing at least one retrieval crawler under the broad binary indicator, the source of the directive was classified into three categories.

56.3%
Explicit — author-set
76 sites. Block is in the site’s own robots.txt. May be intentional or legacy configuration.
39.3%
Indeterminate
53 sites. Origin cannot be confirmed by automated analysis alone; reported as indeterminate rather than resolved.
4.4%
Infrastructure-imposed
6 sites. Block originates from a managed-CDN default the owner may never have consciously set.

Where Australian sites do restrict, the restriction is mostly a deliberate choice rather than a platform default — though a substantial share is indeterminate as to origin and is reported as such.

Sector analysis

Broad-disallow rates by industry

Share of policy-observed domains disallowing at least one retrieval crawler from at least one path — the broad binary indicator, reported for sector comparison. Rates rest on small per-sector denominators (21–44) and should be read with that in mind.

% disallowing ≥1 retrieval crawler (of policy-observed domains per sector)
Real Estate
57.1%
Accounting & Finance
51.6%
Education & Training
50.0%
Technology & SaaS
47.7%
Healthcare
46.9%
Retail & Ecommerce
45.9%
Hospitality & Tourism
38.1%
Legal
36.1%
Building & Trades
31.2%
Professional Services
22.6%

Real Estate (57.1%) and Accounting & Finance (51.6%) show the highest broad-disallow rates; Professional Services (22.6%) and Building & Trades (31.2%) the lowest. These are the broad binary indicator only — necessarily higher than the strict and meaningful measures because they count functional-path disallows.

Crawler panel

Retrieval and training crawlers, reported separately

Block rates across the 21-crawler panel, of 317 policy-observed domains. Retrieval crawlers (Group A) determine whether AI systems can access content to answer queries; training crawlers (Group B) gather data for model training. The two groups are reported separately and never combined.

Group A — retrieval
GPTBot41.0%
ClaudeBot40.7%
anthropic-ai39.7%
Bingbot39.7%
DuckAssistBot39.7%
ChatGPT-User39.4%
Perplexity-User39.1%
OAI-SearchBot38.8%
Googlebot38.5%
PerplexityBot38.2%
Group B — training
CCBot42.0%
Bytespider41.6%
Amazonbot41.3%
Applebot-Extended41.0%
meta-externalagent41.0%
Google-Extended40.7%
FacebookBot40.1%

Group A retrieval-crawler block rates fall in a narrow band (38.2–41.0%): restriction is largely applied across retrieval crawlers as a group rather than targeted at individual operators.

Limitations

What this study does and does not claim

Disclosure & Intellectual Property

Roles. This study is published by the Periodic Table of Digital Authority (PTODA), the publisher and steward of the PTODA research methodology. It was conducted using the PTODA C01 Crawler v1.5.1, a deterministic robots.txt reference instrument, under PTODA C01 Crawler Methodology v1.5. The sample was constructed from named public sources using the published sampling standard. Commercial relationships played no role in domain selection, inclusion, exclusion, analysis, or interpretation. The methodology is fully documented and designed to support independent reproduction using the published specification and frozen datasets. This study publishes aggregate, anonymised findings only. No named individual site results are published.

Attribution chain: Douglas Lord (researcher and author) · Periodic Table of Digital Authority (publisher and methodology steward) · PTODA C01 Crawler v1.5.1 (research instrument) · Digital Dominator Pty Ltd ABN 28 616 931 116 (operating entity).

Intellectual property notice: This study, its methodology, findings, data, and all associated content are the original work of Douglas Lord and the property of Digital Dominator Pty Ltd (ABN 28 616 931 116). The Periodic Table of Digital Authority™ is a coined framework and trade mark pending (TM 2644497). AUTHORITY44™ is a trade mark pending (TM 2643932). All rights reserved.