ansaurus

Question

urllib2.urlopen throws 404 exception for urls that browser opens

Answer 1

+3 A:

Given a URL like:

http://store.ovi.com/#/applications?categoryId=20&amp;fragment=1&amp;page=2

The bit that browsers fetch is just:

http://store.ovi.com/

Everything to the right of that is a ‘fragment identifier’, which is not passed to the server at all (evidently, if you try, it will get confused). Instead, the HTML returned for the / URL will include a load of JavaScript that reads the #... data at the client side and fills in the page content using a bunch of XMLHttpRequests.

Webapps implemented like this are a big old pain to scrape, because you can't just take the HTML content of the main page. Instead you have to either analyse the script to find out where it gets the actual data from, or you have to hook up a real browser in order to execute all the scripts and see what document objects you're left with. They're also typically bad for accessibility and SEO.

Luckily for you this site appears to be putting something in the fragment that's also a valid path. So it looks like you can get the dynamic page data from the URL:

http://store.ovi.com/applications?categoryId=20&amp;fragment=1&amp;page=1

bobince 2010-08-28 01:21:08

Thanks for answering.Opening the url using wget() from the console works with no problem. Why does wget seem to pass the fragment identifier successfully, while urlib2 doesn't?Thanks.

smumbai 2010-08-30 02:00:20

`wget` doesn't pass the fragment identifier, it cuts it off the URL before using it. It's kind of a bug if `urllib2` doesn't.

bobince 2010-08-31 00:53:17

ansaurus

tags:

views:

answers:

urllib2.urlopen throws 404 exception for urls that browser opens

related questions