Showing posts with label poster. Show all posts
Showing posts with label poster. Show all posts

Wednesday, June 22, 2011

SNPBoost: Interaction Analysis of Risk Prediction on GWA Data

Among a few works related to bioinformatics was the poster of Ingrid Braenne et al from Lübeck, Germany, titled SNPboost: Interaction Analysis of Risk Prediction on GWA Data. As commonly known, the human genome consists of DNA, in which the order of the four nucleotides, A, T, C and G, determines the instructions to build structural and functional parts of our bodies. The human genome consists of over three billion nucleotides divided into 23 pairs of chromosomes. Small variations in certain locations of the genome, i.e. genomic loci, make us unique individuals. For example, a single nucleotide variation in a certain loci between individuals is called a single-nucleotide polymorphism, SNP. The human genome contains hundreds of thousands of SNPs. If a SNP is within a gene, the different variants of that gene are called alleles. An individual can be a homozygote or a heterozygote with respect to an allele, implying that a homozygote has the same variant of the gene in both chromosome pairs, whereas heterozygote has different variants in different chromosomes in a chromosome pair. A SNP is usually biallelic, meaning that there are only two high-frequent nucleotide variants of that SNP in the population. Within certain populations, the frequencies of different alleles usually vary, for example, between ethnic groups.

Different variants of a genomic locus, for example a gene, might be responsible for the development of a certain disease. There are certain diseases, which can be associated to a single SNP, such as sickle cell anemia, but most common diseases, i.e. national diseases, such as type I and type II diabetes, myocardial infarction and Chrohn’s disease are so called complex diseases or multi-gene diseases. Complex diseases are caused by several disease-inducing SNPs interacting with each other. SNP profiles of individuals can be obtained using sequencing techniques, and they are routinely analyzed using genome-wide association analysis (GWA). One of the goals in GWA is to associate disease risk to certain SNPs. In addition, the treatment of a disease can be tailored for each individual based on their SNP profiles, leading to personalized medicine. However, the disease diagnostics, or prediction of a disease risk based on SNP profiles, are unfortunately very difficult and uncertain. It is also interesting to determine the functional role of a SNP associated to a certain disease. A SNP can fall into coding regions, i.e. genes, but many of them fall also non-coding regions which makes their functional role harder to depict. Often the interaction of multiple SNPs is not investigated in GWA studies, as GWA only provides the statistical significance of disease association for single SNPs. However, one genetic variant can have only limited effect on risk in complex diseases, and the disease effect may only be visible in the interaction between multiple SNPs. The work of Braenne et al attempts to find solutions for this. In GWA studies, the data created is very high-dimensional and sample size low compared to it. While this causes difficulties to computational analysis, also the interpretation of the results becomes challenging. Therefore, some feature selection is often needed. This issue is also addressed in the work of Braenne et al.

Classifiers, such as support vector machine (SVM), has been successively used, for example, to improve the risk prediction for Type I and type II diabetes from SNP profiles. SVM is a standard benchmark classifying method particularly useful to study multivariate data sets. SVM aims to determine a hyperplane in a higher-dimensional feature space that separates two classes with a maximum separation, i.e. with a maximum margin. By finding a linear separator in a feature space, the classification boundaries are non-linear in the original variable space. In the work of Braenne et al, authors studied the performance of two types of classification tools, SVM and Boosting, to classify healthy and diseased SNP-profiles. Two different versions of SVM were used, namely SVMs with a linear and gaussian kernels. When training a SVM classifier, the authors did feature selection by selecting the subset of SNPs based on their statistical significance of the association with the disease, i.e. their p-values obtained from GWA analysis. Using preselected subset of SNPs might however miss the important interaction effects in the data, so the selection of SNPs for a classifier should be improved somehow.

Boosting method used in the study of Braenne et al is Adaboost, the most popular boosting algorithm. In boosting, the idea is to combine several weak classifiers to obtain one strong classifier. In this work, the set of weak classifiers contained 6 classifiers for each SNP. Each SNP is assumed to be biallelic, and therefore can have three discrete states: AA, AB and BB. AA and BB are homozygotic and AB is a heterozygotic allele. These three states can induce a disease or protect against it, so an individual can be in one of the 2x3=6 possible states; diseased AA, diseased AB, diseased BB, protected AA, protected AB, or protected BB. Weak classifiers try to predict correctly the disease state of an individual based on their genotype and possible state(inducing or protecting) of that SNP. This method is called SNPBoost.

The weak classifiers are combined such that the classification errors of the first single classifier is compensated as good as possible by the second classifier and so forth. These weak classifiers are added one after another in order to gain a set of classifiers that together boost the classification. Selecting classifiers also controls the number of SNP to be included to the classifier. This is claimed to find also possible SNP-SNP interactions between selected SNPs, because if a weak classifier positively interacts with a previous selected one, this weak classifier might be chosen since an interaction might improve the classification.

The data set used in this analysis contained 127370 SNPs, and 2032 individuals with 1222 controls and 810 cases. The data was randomly divided into training and test set both containing 405 individuals. The different classifiers were trained using training data and the classification performance was accessed using test data. Receiver operation characteristics (ROC) curves, i.e. the true positive rate versus the false positive rate were obtained on the test set for each method

Linear SVM (LSVM) with a preselected SNPs and SNPBoost both yielded a peak performance for small number of SNPs. Gaussian SVM (GSVM) yielded a peak in the performance with a slightly larger number of SNPs. The performance decreased with additional SNPs. SNPBoost outperformed the SVM with a linear kernel. Gaussian SVM outperformed linear SVM and SNPBoost, when preselected set of SNPs was chosen to train GSVM. With very large number of SNPs, SNPBoost outperforms all other methods. In addition, with small number of SNPs, the performance of the Gaussian SVM further improved when the subset of SNPs was chosen based on SNPBoost. The better selection of SNPboost compared to LSVM might be due to the fact that the SNPboost algorithm selects SNPs one at a time in order to bit by bit increases the classification and these SNPs might just better fit together.

Braenne et al also investigated the biological interpretation of the results. With SVM and preselected gene sets, 14 SNPs were found within genes. Two of these genes can be linked through an additional gene, meaning that there might be some functional relationships between these genes. When studying the 20 SNPs selected by the SNPboost algorithm, and the their corresponding genes, 9 out of 20 lie within genes and two of the genes can be linked through an additional gene. Both methods revealed one set of three genes, which possible functional relationships, but the relationship of the genes selected by SNPBoost might be an additional one. Future work of is to identify whether the SNPs in these genes increase the disease risk.

Tuesday, June 21, 2011

Nandita Tripathi: Hybrid Parallel Classifiers for Semantic Subspace Learning

Searching large data repositories is an extremely important research problem due to the overwhelming information overload we are facing daily. The poster by Tripathi, Oakes and Wermter presents a hybrid parallel classification approach for searching a large data repository more efficiently. The poster presents an approach to a supervised learning problem: learning to classify text data when labeled training data exists.

Many ways of enhancing the predictive performance of classifiers by using them in a subspace of the original input space have been studied in the recent years. These methods include e.g. the Random Subspace Method (RSM), which divides the original feature space into lower dimensional subspaces randomly and variants of the RSM that use some criteria to select the subspaces instead random assignment.

The novel idea of Tripathi et al. is to use semantic information about the data to optimize the selection of subspaces. After learning the subspaces, a set of classifiers is then used to classify the data with respect to some topics in the new subspaces.

To study the approach, the Reuters data is used. The Reuters data provides multiple levels of topics for its documents. The most broad topics (e.g. education, computers, politics) are inferred as semantic information and they are used for learning the lower dimensional subspaces with a maximum significance based method. After this, multiple classifiers are used within the learnt subspaces to classify the data with respect to more fine-grained topics (e.g. within the topic of education: schools, collage, exams).

Tripathi et al. do experiments using multiple different algorithms (e.g. multiple layer perceptrons, naive bayes classifier, random forest...) as a part of their hybrid architecture. The hybrid architecture both improves classification results and decreases computation times.

Friday, June 17, 2011

Kauppi et al: Face Prediction from fMRI Data during Movie Stimulus: Strategies for Feature Selection

The topic of the poster was to predict from a person's fMRI (functional magnetic resonance imaging) data whether he's seeing a face or faces in a movie or not. In an fMRI test setup, a stimulus, in this case the movie Crash, is presented for the test subject. The test subject's brain activity is measured, resulting in high-dimensional brain activity data that contains complex interactions. In the data, the brain is divided into voxels, i.e. cubes or 3D-pixels.

Similar research had been done before, but there the test subjects were shown a set of movie clips, instead of a whole movie. The authors claim that showing a whole movie results in "more naturalistic" data.

The problem is a classification task with two classes: "face" and "non-face". It was solved using ordinary least squares (OLS) for regression. However, since there was a lot of data, OLS couldn't be used in the conventional way. The prediction was done using only a subset of the features, which were selected using prior information and different methods, resulting in four regression models:

  • Stepwise Regression (SWR)
  • Simulated Annealing (SA)
  • Least Absolute Shrinkage and Selection Operator (LASSO)
  • Least Angle Regression (LARS)

Out of which LASSO and LARS are regulated to be sparse, possibly resulting in less overfitting.

Figure: The best prediction acquired with LARS compared to the (roughly binary) annotation. As can be seen, binary prediction (1 when > 0.5) would match the annotation well. Also locations of 6 features in three bain regions visualized.

Human brain is divided into different regions with different tasks. This study provided a natural way (at least for a computer scientist) to find out which regions are associated with face recognition, and thus, can be used in the prediction. In their paper, the authors say, "our results support the view that face detection is distributed across the visual cortex, albeit the fusiform cortex has a strong influence on face detection."