Very basic question about Hadoop and compressed input files

views:

474

answers:

+4 Q:

Very basic question about Hadoop and compressed input files

I have started to look into Hadoop. If my understanding is right i could process a very big file and it would get split over different nodes, however if the file is compressed then the file could not be split and wold need to be processed by a single node (effectively destroying the advantage of running a mapreduce ver a cluster of parallel machines).

My question is, assuming the above is correct, is it possible to split a large file manually in fixed-size chunks, or daily chunks, compress them and then pass a list of compressed input files to perform a mapreduce?

yes, you could have one large compressed file, or multiple compressed files (multiple files specified with -files or the api).

TextInputFormat and descendants should automatically handle .gz compressed files. you can also implement your own InputFormat (which will split the input file into chunks for processing) and RecordReader (which extract one record at a time from the chunk)

another alternative for generic copmression might be to use a compressed file system (such as ext3 with the compression patch, zfs, compFUSEd, or FuseCompress...)

jspcal 2010-01-16 20:58:12

Thanks that sound great.

Luis Sisamon 2010-01-17 06:20:12

You can use bz2 as your compress codec, and this format also can been split.

zjffdu 2010-01-22 09:15:33

+1 A:

Consider using LZO compression. It's splittable. That means a big .lzo file can be processed by many mappers. Bzip2 can do that, but it's slow.

Cloudera had an introduction about it. For MapReduce, LZO sounds a good balance between compression ratio and compress/decompress speed.

Victor 2010-01-30 09:42:09

ansaurus

tags:

views:

answers:

Very basic question about Hadoop and compressed input files

related questions