ansaurus

Question

k-means clustering implementation in python, running out of memory

Answer 1

A:

Your centroids does not need to be an actual list.

You never appear to reference anything other than centroids[i][m]. If you only want centroids[i], then perhaps it doesn't need to be a list; a simple dictionary would probably do.

    centroids = defaultdict(float)

    #  Move the centroids to the average of their members
    for i in range(k):
        len_best = len(bestmatches[i])

        if len_best > 0:             
            items = set.union(*[set(prefs[u].keys()) for u in bestmatches[i]])

            for user_id in bestmatches[i]:
                row = prefs[user_id]
                for m in items:
                    if row[m] > 0.0: centroids[m]+=(row[m]/len_best)

May work better.

S.Lott 2009-08-05 17:05:27

Answer 2

+4 A:

Not all these observations are directly relevant to your issues as expressed, but..:

a. why are the key in prefs, as shown, longs? unless you have billions of users, simple ints will be fine and save you a little memory.

b. your code:

centroids = [prefs[random.choice(users)] for i in range(k)]

can give you repeats (two identical centroids), which in turn would not make the K-means algorithm happy. Just use the faster and more solid

centroids = [prefs[u] for random.sample(users, k)]

c. in your code as posted you're calling a function simple_pearson which you never define anywhere; I assume you mean to call sim_func, but it's really hard to help on different issues while at the same time having to guess how the code you posted differs from any code that might actually be working

d. one more indication that this posted code may be different from anything that might actually work: you set bestmatch=(0,0) but then test with if d < bestmatch[1]: -- how is the test ever going to succeed? is the distance function returning negative values?

e. the point of a defaultdict is that just accessing row[m] magically adds an item to row at index m (with the value obtained by calling the defaultdict's factory, here 0.0). That item will then take up memory forevermore. You absolutely DON'T need this behavior, and therefore your code:

  row = prefs[user_id]                    
  for m in items:
      if row[m] > 0.0: centroids[i][m]+=(row[m]/len_best)

is wasting huge amount of memory, making prefs into a dense matrix (mostly full of 0.0 values) from the sparse one it used to be. If you code instead

  row = prefs[user_id]                    
  for m in row:
      centroids[i][m]+=(row[m]/len_best)

there will be no growth in row and therefore in prefs because you're looping over the keys that row already has.

There may be many other such issues, major like the last one or minor ones -- as an example of the latter,

f. don't divide a bazillion times by len_best: compute its inverse one outside the loop and multiply by that inverse -- also you don't need to do that multiplication inside the loop, you can do it at the end in a separate since it's the same value that's multiplying every item -- this saves no memory but avoids wantonly wasting CPU time;-). OK, these are two minor issues, I guess, not just one;-).

As I mentioned there may be many others, but with the density of issues already shown by these six (or seven), plus the separate suggestion already advanced by S.Lott (which I think would not fix your main out-of-memory problem, since his code still addressing the row defaultdict by too many keys it doesn't contain), I think it wouldn't be very productive to keep looking for even more -- maybe start by fixing these ones and if problems persist post a separate question about those...?

Alex Martelli 2009-08-05 17:23:39

Thanks for your input, I'll update the original question with a link to simple_pearson (which I've pasted elsewhere to avoid clutter here). The sim_func in the method definition is a remnant of older code.

Andrew Ingram 2009-08-05 19:11:52

These hints seem to have done the trick, it's managing to reach the second iteration and beyond. Iterations past the first one are very slow though (about 10 minutes each), I'll let it run overnight and see what happens. Once I've got my centroids I don't imagine I'll have to recalculate them very often anyway.

Andrew Ingram 2009-08-05 21:47:00

ansaurus

tags:

views:

answers:

k-means clustering implementation in python, running out of memory

Note: updates/solutions at the bottom of this question

Update: Final algorithms

related questions