Draft — wording pending Legal review. This page describes the crawler's behaviour as built; the policy language is not final and has not been approved for publication.

RoktCrawl crawler policy

This policy states how Rokt's storefront crawler, RoktCrawl, identifies itself, what it accesses, how it limits its requests, what it retains, and how a site operator can opt out.

1. Scope

Rokt is an e-commerce technology company. We work with online retailers and other businesses to present relevant offers to their customers at the moment of purchase. RoktCrawl is operated by Rokt and visits the public storefronts of businesses that work with us, whether as a partner running our placement or as an advertiser whose campaigns we serve. We crawl a site because understanding how its pages are structured — its search results, product pages and category pages — is what lets our systems classify pages and better serve offers, without you having to send us that information yourself. RoktCrawl is not a general web crawler and does not index the web.

2. Identification

Every request carries the token RoktCrawl/0.1 (+https://www.rokt.com) appended to a standard browser user-agent string. The token names Rokt as the operator and links to this site. The browser and platform named in the rest of the string are the ones actually in use; the crawler does not disguise itself as another client.

3. Robots exclusion

The crawler requests robots.txt before any other page and applies its rules strictly. It honours rules addressed to RoktCrawl and the general rules addressed to *. A disallowed address is not requested, even when the site's own search form produced it. If robots.txt cannot be read because the site withholds or blocks it, the crawler treats the site as disallowing everything and requests nothing further. Rules are read afresh on every visit.

4. What is accessed

A visit consists of: the site's robots.txt, its sitemap files, its homepage, and a browser session that submits a short list of generic product words to the site's own search box and browses a sample of the resulting search, product and category pages, including their sorting, filtering and paging controls.

The crawler does not sign in, does not create accounts, does not add items to a cart, does not enter checkout, and does not submit any form other than the search box. Addresses for cart, bag, checkout, order, account and login pages are excluded from browsing. It does not attempt to bypass rate limiting, bot detection or any other access control: a page that answers with a block or a challenge is recorded as blocked and is not retried.

5. Request rate

Requests to a single site are spaced at least 2.5 seconds apart, with up to 1.5 seconds of additional random delay, and the crawler loads one page of a site at a time. Within a page load, the browser retrieves the page's own assets as an ordinary browser would. Each visit covers a bounded number of pages, and scheduled crawls run approximately weekly per region. Requests time out after 30 seconds and follow at most ten redirects.

6. What is retained

The crawler retains, for each page it loads: the address, the response headers, a copy of the page body, a screenshot, the addresses of requests the page made, and the crawler's own log of the visit. This material is stored in Rokt's own systems and used to derive and verify the address patterns described above and to describe the storefront's technical characteristics (platform, search capabilities, structured data). Retention period: to be set.

The crawler has no account on any site and sees only what an anonymous visitor sees. It does not seek out personal data, and when it later reads search terms out of addresses, values that appear to be a person's details are withheld rather than recorded.

7. Source addresses

The crawler operates only from the fixed network addresses below, so that a site operator can verify that a request identifying itself as RoktCrawl originated from Rokt. A request carrying the RoktCrawl token from any other address did not come from Rokt. The staging column lists the addresses used by Rokt's pre-production environment, which visits the same sites in small volumes.

RegionProductionStaging
US West (Oregon)16.145.117.88 32.186.91.66 184.33.228.452.38.96.62
US East (N. Virginia)3.212.239.12 34.195.230.14 100.56.59.133.93.202.206
Europe (Ireland)54.73.34.189 3.248.25.71 34.242.239.954.220.78.42
Asia Pacific (Sydney)13.54.60.54 32.236.51.124 13.211.111.19713.54.70.40

8. Opting out

A site operator may exclude the crawler at any time with a robots.txt rule:

User-agent: RoktCrawl
Disallow: /

The exclusion takes effect at the next visit. An operator may also ask Rokt to exclude a site directly using the contact below; such requests are honoured without requiring a robots rule.

9. Contact

Questions, exclusion requests and reports about the crawler's behaviour go to crawler@rokt.com. Include the site's hostname and, where available, the time and source address of the requests concerned.

10. Changes to this policy

If the crawler's identification, behaviour or source addresses change, this page is updated before the change takes effect.