Our crawler
What isitsellingme/1.0 is, what it fetches, how often, and how to stop it.
What you are looking at
isitsellingme/1.0 (+https://isitsellingme.com/bot)
If that is in your logs, it is us. This site grades how companies handle a refusal of advertising data collection, and every grade quotes the company's own published policy, verbatim and dated. The crawler exists to fetch and store those policy pages so that a quotation can be checked against a document rather than taken on trust.
What it fetches, and how often
- 32 URLs in total, across every company on the site — privacy policies, cookie notices and advertising-preference pages. Public documents, all of them.
- One pass a day, shortly after 06:15 UTC. That is one request per URL, not a crawl: it never follows links, never discovers new pages, and never touches anything but the addresses on its list.
- No logins, no forms, no search. It reads what you serve to anyone.
- A page is stored only when its content hash changes, but it isfetched every day either way, so the record can say "we looked, and it was the same" rather than leaving a gap.
It sometimes sends a browser User-Agent
We ask under our own name first. If a host refuses that, the crawler tries once more with an ordinary Chrome User-Agent, and if that also fails it looks for a copy in the Internet Archive instead.
That is worth saying out loud rather than leaving for someone to discover. Several companies on this site serve their privacy policy to a browser and refuse it to anything identifying itself as automated — a document they are legally required to publish, withheld from the only kind of reader that would check it systematically. We are not willing to let that be the reason a company goes ungraded.
Every attempt is recorded either way, under its own name, so the archive shows per host whether an identified crawler was allowed to read the policy. If you would rather we always identified ourselves to you, say so and we will pin your host to the honest User-Agent.
It does not consult robots.txt
Most crawlers should. This one fetches a fixed list of 32 public legal documents once a day, and honouring a Disallow would let a company opt out of having its published policy quoted accurately — which is the precise behaviour this site exists to document.
That is a deliberate position, not an oversight, and you are entitled to disagree with it. Blocking us by User-Agent or IP works, we will not route around it, and the entry will say that the document could not be retrieved.
How to stop it, or reach a person
- Block it. Deny on the User-Agent above. One request a day will stop.
- Tell us the page moved. The commonest problem is a stale URL on our side, which wastes your bandwidth and gets us the wrong document. We would rather fix it.
- Dispute what we quoted. Every company gets a right of reply, published verbatim beside its entry, uncut. See the methodology for how grades are derived, and corrections for what we have got wrong.
Contact: the about page has the address, and security.txt covers anything that needs a security contact rather than an editorial one.