Showing posts with label feature selection. Show all posts
Showing posts with label feature selection. Show all posts

Tuesday, June 21, 2011

Schaffernicht et al. : Weighted Mutual Information for Feature Selection


Evening of first day of ICANN conference i.e. 14th of June was reserved for the poster presentations. Out of papers accepted for publication in ICANN’11 proceedings; around 50 of them were provided with the opportunity to present their work orally while remaining (around 60) were provided with the opportunity to present their work as posters. Poster session was held in T-Building of Computer Science department in Aalto University of Science while the regular conference took place in Dipoli Congress Center. There was enthusiastic participation of both poster presenters and other regular attendants of the conference. It was heartening to find that some of the oral presenters were also presenting their work in the form of posters. It made some sense because poster sessions provides one of the best opportunities to discuss own research with fellow researchers and renowned professors in the field which provides new perspective and sometimes new dimension to research.

Out of many posters; I decided to write about this poster by Erik Schaffernicht and Horst-Michael Gross; the topic of which was Weighted Mutual Information for Feature Selection. There is no denying the importance of feature selection in learning algorithms. Hence, the research area is never saturated in this field and have opening for new ideas and methods. In this paper; the authors provide a simple trick to include only the relevant features and at the same time avoid redundant features. Similar to other wrapper methods; they individually train the classifier on the entire features. However, they determine the next feature to be included by their accuracy on the misclassified samples rather than the entire data samples. They provide weights to the samples and select the features which maximize the mutual information.

This idea is similar to the well known AdaBoost algorithm. The misclassified samples in the first round are given higher weights in the second round. This makes some sense because correctly classified samples are easily explained by the selected subset of features at any time instant and the crux of the problem is to find those features that better classify the misclassified samples. They experimented this methods in different datasets from UCI machine learning repository and also on some artificial datasets. Although it was not mentioned in the paper or displayed in poster but I enquired with them if they are using that in some real world datasets for current project and they informed me that they had deployed that in a control systems where data dimension is in thousands.

Overall, a simple but very effective and intelligent trick to select features. It achieves computational efficiency by reducing training cycles significantly and also selects the set of best discriminating features.

Friday, June 17, 2011

Kauppi et al: Face Prediction from fMRI Data during Movie Stimulus: Strategies for Feature Selection

The topic of the poster was to predict from a person's fMRI (functional magnetic resonance imaging) data whether he's seeing a face or faces in a movie or not. In an fMRI test setup, a stimulus, in this case the movie Crash, is presented for the test subject. The test subject's brain activity is measured, resulting in high-dimensional brain activity data that contains complex interactions. In the data, the brain is divided into voxels, i.e. cubes or 3D-pixels.

Similar research had been done before, but there the test subjects were shown a set of movie clips, instead of a whole movie. The authors claim that showing a whole movie results in "more naturalistic" data.

The problem is a classification task with two classes: "face" and "non-face". It was solved using ordinary least squares (OLS) for regression. However, since there was a lot of data, OLS couldn't be used in the conventional way. The prediction was done using only a subset of the features, which were selected using prior information and different methods, resulting in four regression models:

  • Stepwise Regression (SWR)
  • Simulated Annealing (SA)
  • Least Absolute Shrinkage and Selection Operator (LASSO)
  • Least Angle Regression (LARS)

Out of which LASSO and LARS are regulated to be sparse, possibly resulting in less overfitting.

Figure: The best prediction acquired with LARS compared to the (roughly binary) annotation. As can be seen, binary prediction (1 when > 0.5) would match the annotation well. Also locations of 6 features in three bain regions visualized.

Human brain is divided into different regions with different tasks. This study provided a natural way (at least for a computer scientist) to find out which regions are associated with face recognition, and thus, can be used in the prediction. In their paper, the authors say, "our results support the view that face detection is distributed across the visual cortex, albeit the fusiform cortex has a strong influence on face detection."