CS231n笔记 Lecture 11, Detection and Segmentation

Other Computer Vision Tasks

Semantic Segmentation. Pixel level, don't care about instances.
Classification + Localization. Single object.
Object Detection. Multiple object.
Instance Segmentation. Multiple object.

Semantic Segmentation

Simple idea: sliding window, crop across the whole image, and ask what the center pixel is. Expensive.

Fully Convoltional (Naive) : let the network to learning all the pixels at once, keep the spacial size, convolutions at original image resolution, expensive.

Fully convolutional: Design network as a bunch of convolutional layers, with downsampling and upsampling inside the network!

Downsampling: Pooling, strided convolution
Upsampling: Unpooling (nearest neighbor, bed of nails, max unpooling in symetrical NN), Transpose convolution (multiply the filter by the pixels on the input, use stride and pad to impose the value on the output).