ansaurus

Question

passing text through a dictionary in Python

Answer 1

+1 A:

If the dictionary fites in memory, use a Python set:

ok_words = set(["a", "b", "c", "e"])

def filter_words(words):
    return [word for word in words if word in ok_words]

If it doesn't fit in memory, you can use shelve

lazy1 2010-10-21 23:56:38

I believe last string must be `return [word for word in words **if** word in ok_words]` ?

Andrei 2010-10-22 00:45:39

I updated the typo.

Nick Presta 2010-10-22 00:54:38

Sets are not the same thing as dictionaries (though the implementation is similar).

intuited 2010-10-22 01:46:17

Thank you. This is the "pythonic" way I was looking for. I'll have to look more into shelve though.

KyleP 2010-10-22 15:42:49

Andrei: You're right, my bad (note to self: *run* the code before posting)

lazy1 2010-10-22 20:48:10

Answer 2

A:

The structure you try to create is known as Inverted Index. Here you can find some general information about it and snippets from Heaps and Mills's implementation. Unfortunately, I wasn't able to find it's source, as well as any other efficient implementation. (Please leave comment if you will find any.)

If you haven't a goal to create a library in pure Python, you can use PyLucene - Python extension for accessing Lucene, which is in it's turn very powerful search engine in Java. Lucene implements inverted index and can easily provide you information on word frequency. It also supports wide range of analyzers (parsers + stemmers) for a dozen of languages.
(Also note, that Lucene already has it's own Similarity measure class.)

Some words about similarity and Vector Space Models. It is very powerful abstraction, but your implementation suffers several disadvantages. With a growth of number of documents in your index your co-occurrence matrix will became to big to fit in memory, and searching in it will take a long time. To stop this effect dimension reduction is used. In methods like LSA this is done by Singular Value Decomposition. Also pay attention to such techniques as PLSA, which uses probabilistic theory, and Random Indexing, which is the only incremental (and so the only appropriate for the large indexes) VSM method.

Andrei 2010-10-22 00:35:10

Since this topic is not about VSM, I won't give more info about it here, but if you need it, please create new topic and post a comment with it here.

Andrei 2010-10-22 00:43:22

Thank you for the links.

KyleP 2010-10-22 15:38:53

ansaurus

tags:

views:

answers:

passing text through a dictionary in Python

related questions