ansaurus

Question

Answer 1

+4 A:

I used to use simple html dom exclusively until some bright SO'ers showed me the light hallelujah.

Just use the built in DOM functions. They are written in C and part of the PHP core. They are faster more efficient than any 3rd party solution. With firebug, getting an XPath query is muey simple. This simple change has made my php based scrapers run faster, while saving my precious time.

My scrapers used to take ~ 60 megabytes to scrape 10 sites asyncronously with curl. That was even with the simple html dom memory fix you mentioned.

Now my php processes never go above 8 megabytes.

Highly recommended.

EDIT

Okay I did some benchmarks. Built in dom is at least an order of magnitude faster.

Built in php DOM: 0.007061
Simple html  DOM: 0.117781

<?
include("../lib/simple_html_dom.php");

$html = file_get_contents("http://stackoverflow.com/search?q=favorite+programmer+cartoon");
$data['dom'] = $data['simple_dom'] = array();

$timer_start = microtime(true);

$dom = new DOMDocument();
@$dom->loadHTML($html);
$x = new DOMXPath($dom); 

foreach($x->query("//a") as $node) 
{
     $data['dom'][] = $node->getAttribute("href");
}

foreach($x->query("//img") as $node) 
{
     $data['dom'][] = $node->getAttribute("src");
}

foreach($x->query("//input") as $node) 
{
     $data['dom'][] = $node->getAttribute("name");
}

$dom_time =  microtime(true) - $timer_start;

echo "built in php DOM : $dom_time\n";

$timer_start = microtime(true);
$simple_dom = new simple_html_dom();
$simple_dom->load($html);
foreach( $simple_dom->find("a") as $node)
{
   $data['simple_dom'][] = $node->href;
}

foreach( $simple_dom->find("img") as $node)
{
   $data['simple_dom'][] = $node->src;
}

foreach( $simple_dom->find("input") as $node)
{
   $data['simple_dom'][] = $node->name;
}
$simple_dom_time =  microtime(true) - $timer_start;

echo "simple html  DOM : $simple_dom_time\n";

Byron Whitlock 2010-08-30 19:32:25

this woun't work for invalid markup. How much faster is this versus simple dom?

Quamis 2010-08-30 19:44:51

This **does** work for invalid markup. I don't have benchmarks but it is at least an order of magnitude faster. On large pages, simple html dom would take 1-2 seconds. The built in DOM does it in the blink of an eye. I've written many scrapers with this and I would never use simple html dom for anything ever again.

Byron Whitlock 2010-08-30 19:47:57

@Quamis Notice the @ in front of loadHtml(). With that removed you will see a ton of warnings from invalid html being coerced into the dom tree. Works for browsers, works for php too ;)

Byron Whitlock 2010-08-30 19:49:35

you're right, its way faster and it loads invalid html, just re-tested it now

Quamis 2010-08-30 20:23:46

You can find this benchmark at http://whitlock.ath.cx/FastCrawl/benchmark.php

Byron Whitlock 2010-08-30 20:37:28

i edited your answer to use timestamp(true) instead of simply timestamp().

Quamis 2010-08-30 21:30:28

ansaurus

tags:

views:

answers:

html scraping and css queries

related questions