RoktCrawl: Rokt's storefront crawler
If RoktCrawl appears in your server logs, this page explains what it is,
what it does on your site, and how to reach us or ask it to stop.
Who we are
Rokt is an e-commerce technology company. We work with online retailers and other businesses to present relevant offers to their customers at the moment of a purchase. RoktCrawl is operated by Rokt and visits the public storefronts of businesses that work with us.
What the crawler does
RoktCrawl learns how a storefront's web addresses are structured: which URLs are search results, which are product pages and which are category pages, and how a search term or product identifier is written into the address. Rokt's systems use that knowledge to recognise the kind of page a shopper is on from its address alone, without inspecting the page itself.
A visit fetches the site's robots.txt, its sitemap and its homepage, then opens a real
browser session that types a handful of ordinary search terms into the site's own search box and
browses the results: sorting, filtering, paging, and a sample of product and category pages. The
crawler records the addresses it saw, a copy of each page it loaded and a screenshot, so the
patterns can be derived without visiting again.
How to recognise it
Every request the crawler makes carries this token at the end of an otherwise standard browser user-agent string:
RoktCrawl/0.1 (+https://www.rokt.com)
A full user-agent string from the crawler looks like this (the Chrome version follows the browser build in use):
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Safari/537.36 RoktCrawl/0.1 (+https://www.rokt.com)
The token names the operator and links here. The rest of the string is the real browser the crawler drives; nothing about the platform or browser is disguised.
What it does not do
- It never logs in or creates an account. Pages behind a sign-in are never requested.
- It never enters the cart or checkout. Links to cart, bag, checkout, account, order and login pages are excluded from browsing.
- It submits no form other than the search box, and types nothing into it but a short list of generic product words.
- It does not evade access controls. A page that answers with a block or a challenge is
recorded as blocked and not retried; a site whose
robots.txtis withheld is treated as disallowing everything. - It collects no personal data. It has no account on your site and sees only what an anonymous visitor sees. Search terms it extracts from addresses later are screened so that anything that looks like a person's details is withheld.
robots.txt
The crawler reads robots.txt before anything else and obeys it strictly. A URL that
robots.txt disallows is not fetched: if the site's search form produces a disallowed address, the
address is noted but the page is never loaded. The crawler matches rules addressed to the token
RoktCrawl as well as the general * rules. To keep it off your site entirely, add:
User-agent: RoktCrawl
Disallow: /
The rule takes effect at the crawler's next visit; there is no cached copy of your robots file between visits.
Request rate
Page loads to one site are spaced at least 2.5 seconds apart, plus up to 1.5 seconds of random extra delay, and the crawler visits one page of a site at a time. Within a page load, the browser fetches the page's own images, styles and scripts as any browser would. A site's visit covers a small number of pages, and scheduled crawls run about once a week per region.
Egress addresses
The crawler leaves Rokt's network only through the fixed addresses below, so a request claiming to be RoktCrawl can be verified against them. The full list, with the policy that governs it, is on the crawler policy page.
| Region | Production | Staging |
|---|---|---|
| US West (Oregon) | 16.145.117.88 32.186.91.66 184.33.228.4 | 52.38.96.62 |
| US East (N. Virginia) | 3.212.239.12 34.195.230.14 100.56.59.13 | 3.93.202.206 |
| Europe (Ireland) | 54.73.34.189 3.248.25.71 34.242.239.9 | 54.220.78.42 |
| Asia Pacific (Sydney) | 13.54.60.54 32.236.51.124 13.211.111.197 | 13.54.70.40 |
Contact and opting out
The quickest way to stop the crawler is the robots.txt rule above. To ask us to exclude
your site, to report a problem with the crawler's behaviour on your site, or to verify that a request
came from us, contact us at:
Please include your site's hostname and, if you have it, the time and source address of the requests you are asking about.
The crawler policy states these commitments formally.