ansaurus

Question

Answer 1

+1 A:

EDIT: added some details about explain().

Some general introduction: The Lucene Highlighter is meant to find text snippets from a hit document, and to highlight tokens matching the query.

Therefore, The TokenStream parameter is used to break the hit text into tokens. The highlighter's scorer then scores each token, in order to score fragments and choose snippets and tokens to be highlighted.
I believe you are doing it wrong. If all you want to do is understand which query terms were matched in the document, you should use the explain() method. Basically, after you have instantiated a searcher, use:

Explanation expl = searcher.explain(query, docId);

String asText = expl.toString();

String asHtml = expl.toHtml();

docId is the raw document id from the search results.

Only if you do need the snippets and/or highlights, you should use the Highlighter. If you still want to use the highlighter, follow Nicholas Hrychan's advice. Beware, though, as he describes the Lucene 2.4.1 API - If you use a more advanced version, you should use "QueryScorer" where he says "SpanScorer" .

Yuval F 2010-03-10 12:07:58

I have not understood the explain method. It returns an Explanation object, what function is required hereon to get the matched query terms. I am not satisfied with the documentation of Lucene.

iamrohitbanga 2010-03-10 16:21:40

Please see my edit about explain().

Yuval F 2010-03-11 07:21:39

ok cool. what about getDetail() method.

iamrohitbanga 2010-03-11 07:36:19

The Lucene Explanation has a recursive structure. toString() and toHtml() give you the full explanation tree. getDetails() gives you a subtree of the tree at a time. I would try looking at the full tree first, and only if this is too complicated go to the subtrees.

Yuval F 2010-03-11 10:03:24

this is what i get when i search for skin. the 0th document hit produces this with explain().`0.0 = (NON-MATCH) sum of:`how do i interpret it?

iamrohitbanga 2010-03-12 13:04:22

System.out.println(searcher.explain(query, i).toString());

iamrohitbanga 2010-03-12 13:15:00

i is zero in the previous function call

iamrohitbanga 2010-03-12 13:24:36

Here's a fuller version: http://stackoverflow.com/questions/1742124/different-lucene-search-results-using-different-search-space-size

Yuval F 2010-03-12 15:53:31

OK thank you finally got it working. Now i am using fuzzy query matching. so if the document is a hit, explain would tell me the word as it occurs in the document and not in the query. how to achieve this?

iamrohitbanga 2010-03-12 16:26:26

ansaurus

tags:

views:

answers:

using hit highlighter in lucene

related questions