Showing posts with label Deep learning. Show all posts
Showing posts with label Deep learning. Show all posts

Wednesday, June 22, 2011

Geoffrey Hinton: Learning structural descriptions of objects using equivariant capsules

The second plenary talk in ICANN 2011 was given by Prof. Geoffrey Hinton from University of Toronto. The topic was "Learning structural descriptions of objects using equivariant capsules". The accompanied paper in the proceeding is under the name: “Transforming Auto-encoders”. In this talk, he discussed the limitation of the convolutional neural network, and proposed a new way of learning invariant features under a new neural network framework.



The human brain does not need to go through a step of rotation to recognize an object. This is proven by a test where the task is to recognize objects positioned in arbitrary angles versus the task of imaginatively rotating the same object. However, in several recently popular computer vision algorithms, this rule is violated.

In most popular computer vision research, people use explicitly designed operators to extract the invariant features from images. These operators, according to Prof. Hinton, turn out to be misleading and not efficient. For instance, using convolutional neural network, one will try to learn the invariant features in different part of the images, and discard the spatial relationship between them. This will not work in a higher level features where we need to do, for instance, face identity analysis, which requires extremely strong spatial relationship between mouth and eyes.



Prof. Hinton arguess that the convolutional network way of representing the invariant features, where only some scalar output is used to represent the presence of the feature, is not capable of representing highly complex invariant feature sets. Subsampling methods have been proposed to make convolutional neural networks invariant for small changes in the viewing angle of the object. Prof. Hinton argues that it is not correct as the ultimate goal of learning feature should not be viewpoint invariant. Instead, the goal should be Equivariant features where changes in viewpoints lead to corresponding changes in neural networks. Equivariant feature means that the building block of the object features should be rotated correspondingly while the objects are rotated.

Therefore, he developed a new way of learning feature extractors which learn equivariant features through computation on local space called "capsules", and output informative results. These local features are accumulated hierarchically towards a more abstract representation. The network is then trained with images of the same objects when they are slightly shifted and rotated. In this way, each learned capsule is a "generative model". The difference between convolutional neural network and the "capsule method" is that the capsule method considers the spatial relationship of image features carrying spatial position along with the feature presence probability distribution.

This new way of representing the transformation of images has opened a new possibility for training invariant features and Prof. Hinton argues that this approach behaves closer to the way human brain functions and will be more promising one comparing to traditional computer vision methods.

For detailed explanation and demonstration, please see the full paper included in the proceeding of ICANN 2011.

Thursday, June 16, 2011

Heess, N., Le Roux, N. and Winn, J.. Weakly Supervised Learning of Foreground-Background Segmentation using Masked RBMs

This work (Heess et al., 2011) has been presented in ICANN 2011 as a part of the poster session.

The main target of this paper is to show that a generative model based on restricted Boltzmann machines can be used to distinguish a foreground object (an object in interest) and a background image.

The proposed model starts from a layer of image pixels corresponding to a single image with two directed edges going forward to two separate layers that describe a foreground object and a background image, respectively. Then, each of those layers are connected to a separate layer of latent variables with undirected edges forming an restricted Boltzmann machine. While there exists an additional set of binary variables that denotes a mask of the foreground object in the image, and it is connected to the latent variables that were connected to the layer of the foreground object by the undirected edges.

In other words, there are two RBMs that model (1) jointly appearance and a shape of a foreground object (will be denoted as fRBM from now for simplicity) and (2) a background image, and they are conditioned on the original image (will be denoted as bRBM for simplicity).

This approach suggests that when it is possible to have good generative models for two distinct types of images (or in fact, any other kinds of data sets) it will be able to use them for separating a mixed image (in this case, simply foreground + background). Also, considering the depth of the proposed model (a directed layer + an undirected layer), it can be considered as one of the early approaches for applying deep learning to image segmentation tasks, see (Socher et al., 2011) for another possibility.

One important contribution of this approach is that it does not require explicit ground-truth segmentation of training samples to train the model. Instead, the authors initialize bRBM by training it with images that can be considered easily as backgrounds. Intuitively, this method drives fRBM to learn regularities found by the foreground objects in the training samples while background clutters are considered to be already well-modeled by bRBM. This is a neat trick, but they needed some more tricks in learning process in order to overcome some apparent problems such as training samples having regular structure in the background (such as photos of people taken in a single space).

The experimental results are impressive. However, more experiments on some other data might have been useful for readers to understand the value of the proposed model and learning method. The authors provide interesting future research directions such as; replacing RBMs with deep models, including few ground-truth segmentations to make it into semi-supervised learning, and another layer of hidden nodes immediately after the original image layer.

References
Heess, N., Le Roux, N. and Winn, J.. Weakly Supervised Learning of Foreground-Background Segmentation using Masked RBMs. ICANN 2011.
Le Roux, N., Heess, N., Shotton, J., Winn, J.. Learning a Generative Model of Images by Factoring Appearance and Shape.