CRAWLER POLICY / HUMANOIDONLINEMARKETBOT
CRAWLER POLICY

HumanoidOnlineMarketBot

What it is

HumanoidOnline operates an automated market-observation crawler, HumanoidOnlineMarketBot. Its purpose is to maintain evidence-backed information about humanoid robotics products and their commercial availability: which robots exist, their published specifications, prices and whether they can be ordered, together with the source each fact came from.

Every value it reads is recorded as an unverified observation with its source page and a short supporting excerpt. Nothing it reads is published on HumanoidOnline until a person has reviewed it.

How to recognise it

Every request carries exactly this user agent, and never a browser's:

HumanoidOnlineMarketBot/0.1 (+https://humanoidonline.com/crawler-policy)

The robots.txt product token is HumanoidOnlineMarketBot.

Request rate

  • One request at a time per site, at least 2 seconds apart. If your robots.txt sets a longer Crawl-delay, the longer delay is used.
  • At most 50 product or announcement pages per site in one run, plus the few fixed listing pages they are found on. Links are followed one level only.
  • Runs are started manually by a named member of our team. Unchanged pages are requested conditionally (If-None-Match / If-Modified-Since).
  • Retries are limited to two, only after network errors or server errors (5xx).

robots.txt is respected

robots.txt is read at the start of every run and checked for every page and every redirect before it is requested. If robots.txt disallows a page, that page is not requested. If it disallows our crawler, the run stops and the site is disabled until a person reviews it again. A 401 or 403 on robots.txt is treated as a complete disallow. A 401, 403 or 429 on any page stops the run.

What it does not do

  • It does not run JavaScript, use a headless browser, log in, or solve CAPTCHAs.
  • It does not try to get around access controls, rate limits or blocks.
  • It does not keep copies of your pages. It keeps hashes, short excerpts (at most 1,000 characters) as evidence, and links to images, not the images themselves.

Opting out or slowing it down

To exclude HumanoidOnlineMarketBot from your whole site, add to your robots.txt:

User-agent: HumanoidOnlineMarketBot
Disallow: /

To make it visit less often, set a longer crawl delay (in seconds):

User-agent: HumanoidOnlineMarketBot
Crawl-delay: 30

Both take effect from the next run, because robots.txt is re-read every time. You can also disallow only some paths.

Contact

Site owners can reach us through the contact form at https://humanoidonline.com/contact. Choose “Crawler / robots” or “Crawler opt-out / rate-limit request” and include your site's domain. Use it to:

  • ask questions about the crawler;
  • request a reduced crawl frequency;
  • report a problem caused by the crawler;
  • request that your site be excluded.

An exclusion or rate request is applied to our crawler configuration for your site. Your robots.txt rules are honoured whether or not you contact us.