ansaurus

Question

Answer 1

+2 A:

Any visitor whose entry page is /robots.txt is probably a bot.

Sparr 2009-02-13 02:06:57

Or, to be less strict, a visitor who requests robots.txt at all is probably a bot, although there are a few firefox plugins that grab it while a human is browsing.

Sparr 2009-02-17 21:04:36

Any bot that goes there is probably a well-behaved, respectable bot, the kind you might want visiting your site :-)

Hightechrider 2010-04-23 21:15:33

Answer 2

+3 A:

You can find a very thorough database of data on known "good" web crawlers in the robotstxt.org Robots Database. Utilizing this data would be far more effective than just matching bot in the user-agent.

Sparr 2009-02-13 02:12:00

Answer 3

+3 A:

One suggestion is to create an empty anchor on your page that only a bot would follow. Normal users wouldn't see the link, leaving spiders and bots to follow. For example, an empty anchor tag that points to a subfolder would record a get request in your logs...

<a href="dontfollowme.aspx"></a>

Many people use this method while running a HoneyPot to catch malicious bots that aren't following the robots.txt file. I use the empty anchor method in an ASP.NET honeypot solution I wrote to trap and block those creepy crawlers...

Dscoduc 2009-02-13 02:18:43

Just out of curiosity, this made me wonder if that might mess with accessibility. Like if someone could accidentally select that anchor using the Tab key and then hit Return to click it after all. Well, apparently not (see http://jsbin.com/efipa/ for a quick test), but of course I've only tested with a normal browser.

Arjan 2009-10-08 16:40:01

Need to be a little bit careful with techniques like this that you don't get your site blacklisted for using blackhat SEO techniques.

Hightechrider 2010-04-23 21:14:00

Answer 4

+1 A:

Something quick and dirty like this might be a good start:

return if request.user_agent =~ /googlebot|msnbot|baidu|curl|wget|Mediapartners-Google|slurp|ia_archiver|Gigabot|libwww-perl|lwp-trivial/i

Note: rails code, but regex is generally applicable.

Brian Armstrong 2010-06-25 22:43:04

ansaurus

tags:

views:

answers:

Detecting honest web crawlers

related questions