ansaurus

Question

Calculating Mutual Information For Selecting a Training Set in Java

Answer 1

A:

2010-01-04 23:02:00

Please see my edit, I'm aware of Weka, but have been unable to find resources on what it can do for Information Gain. Could you be more specific? Thanks for your time!

Grundlefleck 2010-01-04 23:21:14

Weka documentation is not the best; if you have a library specific question, I suggest you try the mailing list.

2010-01-05 04:05:08

Answer 2

+1 A:

I know information gain only in connection with decision trees (DTs), where in the construction of a DT, the split to make on each node is the one which maximizes information gain. DTs are implemented in Weka, so you could probably use that directly, although I don't know if Weka lets you calculate information gain for any particular split underneath a DT node.

Apart from that, if I understand you correctly, I think what you're trying to do is generally referred to as active learning. There, you first need some initial labeled training data which is fed to your machine learning algorithm. Then you have your classifier label a set of unlabeled instances and return confidence values for each of them. Instances with the lowest confidence values are usually the ones which are most informative, so you show these to a human annotator and have him/her label these manually, add them to your training set, retrain your classifier, and do the whole thing over and over again until your classifier has a high enough accuracy or until some other stopping criterion is met. So if this works for you, you could in principle use any ML-algorithm implemented in Weka or any other ML-framework as long as the algorithm you choose is able to return confidence values (in case of Bayesian approaches this would be just probabilities).

With your edited question I think I'm coming to understand what your aiming at. If what you want is calculating MI, then StompChicken's answer and pseudo code couldn't be much clearer in my view. I also think that MI is not what you want and that you're trying to re-invent the wheel.

Let's recapitulate: you would like to train a classifier which can be updated by the user. This is a classic case for active learning. But for that, you need an initial classifier (you could basically just give the user random data to label but I take it this is not an option) and in order to train your initial classifier, you need at least some small amount of labeled training data for supervised learning. However, all you have are unlabeled data. What can you do with these?

Well, you could cluster them into groups of related instances, using one of the standard clustering algorithms provided by Weka or some specific clustering tool like Cluto. If you now take the x most central instances of each cluster (x depending on the number of clusters and the patience of the user), and have the user label it as interesting or not interesting, you can adopt this label for the other instances of that cluster as well (or at least for the central ones). Voila, now you have training data which you can use to train your initial classifier and kick off the active learning process by updating the classifier each time the user marks a new instance as interesting or not. I think what you're trying to achieve by calculating MI is essentially similar but may be just the wrong carriage for your charge.

Not knowing the details of your scenario, I should think that you may not even need any labeled data at all, except if you're interested in the labels themselves. Just cluster your data once, let the user pick an item interesting to him/her from the central members of all clusters and suggest other items from the selected clusters as perhaps being interesting as well. Also suggest some random instances from other clusters here and there, so that if the user selects one of these, you may assume that the corresponding cluster might generally be interesting, too. If there is a contradiction and a user likes some members of a cluster but not some others of the same one, then you try to re-cluster the data into finer-grained groups which discriminate the good from the bad ones. The re-training step could even be avoided by using hierarchical clustering from the start and traveling down the cluster hierarchy at every contradiction user input causes.

ferdystschenko 2010-01-05 09:54:54

"Instances with the lowest confidence values are usually the ones which are most informative" - I don't think this is the case in my scenario. There are a few features for each instance and it is basically the instance which matches the most features from other unlabeled data that is the most informative. Is this such an unusual scenario that it's unlikely to be available in a library?

Grundlefleck 2010-01-05 22:50:48

The term 'informative' is meant a little differently here: you can have an instance with all values set but if your classifier already 'knows' which class it belongs to, adding this instance to your training set will not improve accuracy. On the other hand, if an instance is classfied incorrectly, adding it to your training set may give the learner new information - given that there are enough non-zero features, of course, i.e. an instance which couldn't be classfied due to data sparsity, can in principle not provide new information.

ferdystschenko 2010-01-06 09:26:06

I think I may have confused the meaning of Training Data - I was using it as the data given to a human user to classify, which I guess is wrong. The result is that I think you've given an answer which applies to the scenario *just after* the bit I'm talking about. I'm looking for what examples to give to the human in the very first time, before a learning algorithm is involved. Cheers!

Grundlefleck 2010-01-08 12:42:31

Answer 3

+2 A:

StompChicken 2010-01-08 00:20:24

Ah, I think my poor terminology may have let me down. Say I have a list of 'things', I'll call them 'reports'. A report has several features, but what I'm ultimately interested in is the probability of a true/false value to assign to the report. Is the report then a random variable? I may also have used the term feature incorrectly. If the analogy was that a report was a Java object, would a feature be an instance variable of that object? Hopefully if I get my terminology correct I can make the question better.

Grundlefleck 2010-01-08 11:05:34

This was extremely helpful, but as is always the case, it leads me to another question :-p If I calculate mi for every discrete value of each feature, then retrieve the mi value of each feature of a given example, and sum them, would that make sense as an ordering for all the examples? Thanks! PS. I have edited the question which may clarify what I'm meaning some more.

Grundlefleck 2010-01-08 12:39:02

ansaurus

tags:

views:

answers:

Calculating Mutual Information For Selecting a Training Set in Java

related questions