・Write to us at https://ai-lifeline.org/jp/tousho/ (the form is in Japanese; English text is fine).
・We accept exclusion requests for any reason, without asking why.
The domain is removed from our target list and we keep an internal record that it was removed (we do not publish which domains were removed).
・Published by Pasokon Sensei Co., Ltd. (publisher: Masahiro Suenaga).
1. Monthly robots.txt observation
- ・User-Agent:
ai-lifeline-robots-watch/0.1 (+https://ai-lifeline.org/jp/crawler/; monthly robots.txt observation) - ・What we fetch: only
/robots.txton each site. No article text, no images, no other pages. - ・How many: 26 major Japanese sites (news, magazines/publishers, platforms, reference, e-commerce, government).
- ・How often: once per month, one request per site, at least 300 ms apart, 20 s timeout.
- ・We honour robots.txt: if a site disallows our User-Agent, we do not crawl it.
- ・If we get a 403 or hit bot protection: we do not work around it. We do not retry with a different User-Agent and we do not retry later. We simply record it as "could not verify".
- ・Why: to count how often AI crawlers are disallowed. We publish counts and ratios; we never republish the contents of anyone's robots.txt.
2. Weekly public-price observation
- ・User-Agent:
ai-lifeline-collector/0.1 (+https://ai-lifeline.org/jp/crawler/; weekly public-price watch; 1req/week) - ・What we fetch: only the official pricing pages listed in our registry. We do not follow links or expand to other pages.
- ・How many: 2 pages for GPU hourly rates and 4 pages for inference token prices (6 pages in total).
- ・How often: once a week at most, one request per page, at least 500 ms apart.
- ・We honour robots.txt: each target is recorded in the registry together with the date we checked that robots.txt permits it.
- ・If we get a 403 or hit bot protection: we do not work around it. We record it as "could not verify" and drop that provider from the series.
- ・How we publish: we extract price figures only and typeset them ourselves. We do not reproduce prose, tables or images. Every figure carries its source and the date it was retrieved.
2b. Monthly check of AI-Friendly Declarations (added 2026-09-13; registered sites only)
- ・Targets: only sites that registered themselves under the AI-Friendly Declaration. We never visit a site that did not register.
- ・What we fetch: that site's
/.well-known/ai-friendly.json(the declaration), its/robots.txt, and the one machine-readable data URL named in the declaration. Nothing else. - ・How often: once per month (the same day as the robots.txt observation in section 1) plus once at registration. One request per file, at least 300 ms apart, 20 s timeout.
- ・User-Agent:
ai-lifeline-robots-watch/0.1 (+https://ai-lifeline.org/declaration/; ai-friendly declaration check, monthly, opt-in)for robots.txt;ai-lifeline-collector/0.1 (+https://ai-lifeline.org/declaration/; ai-friendly declaration check, monthly, opt-in)for the declaration and the data URL. - ・If we get a 403 or hit bot protection: we do not work around it. We record "not verified" (which is not a negative finding).
- ・How to stop it: ask for removal through the letter box; we remove the site for any reason.
3. What we never do
- ・Work around robots.txt or bot protection (spoofed User-Agents, headless browsers, rotating IPs).
- ・Crawl broadly. Every target is enumerated as an explicit URL in a registry file.
- ・Redistribute or re-host the pages we fetch.
- ・Collect training data. We do not train models.
4. What happens when someone objects
Stop collecting → stop publishing → look for an alternative → record and report, in that order. We stop first and think afterwards.
5. Crawling this site (the other direction)
Crawling for indexing and for AI training is welcome here. See /robots.txt, /llms.txt and AIの方へ (for AI readers).
日本語版:クローラーポリシー
Last updated: 2026-08-26