ansaurus

Question

How to count term frequency for set of documents?

Answer 1

A:

I don't know Lucene, however; your naive implementation will scale, provided you don't read the entire document into memory at one time (i.e use an on-line parser). English text is about 83% redundant so your biggest document will have a map with 85000 entries in it. Use one map per thread (and one thread per file, pooled obviouly) and you will scale just fine.

Update: If your term list does not change frequently; you might try building a search tree out of the characters in your term list, or building a perfect hash function (http://www.gnu.org/software/gperf/) to speed up file parsing (mapping from search terms to target strings). Probably just a big HashMap would perform about as well.

Justin 2010-05-27 19:20:00

Answer 2

A:

See if this helps

Mikos 2010-05-27 20:13:20

thanks, but with this solution i get the overall frequency and not just the frequency for a subset of documents.

ManBugra 2010-05-27 20:26:07

Perhaps you should consider creating a temp index for the sub-set of documents. This might be a hack approach, but you should get the all the power that Lucene provides.

Mikos 2010-05-31 12:57:56

Answer 3

+1 A:

Go here: http://lucene.apache.org/java/3_0_1/api/core/index.html and check this method

org.apache.lucene.index.IndexReader.getTermFreqVectors(int docno);

you will have to know the document id. This is an internal lucene id and it usually changes on every index update (that has deletes :-)).

I believe there is a similar method for lucene 2.x.x

Toader Mihai Claudiu 2010-05-27 20:30:46

Answer 4

A:

See this similar SO question which has a link to some example code.

mindas 2010-05-27 21:02:27

ansaurus

tags:

views:

answers:

How to count term frequency for set of documents?

related questions