Machine Learning for Signal Processing

E9 205 • Fall 2018

Announcements

Class Location

Moved to B-303 EE (Note the Change).

Course Enrollment Form

Please fill this up in case you are interested in audit/credit.

Fifth Assignment

Posted. Due on 23-11-2018.

HW5

Second Mid-Term Exam Date and Time

Nov. 9, 4:00 PM, B308. Open Book, Open Notes.

Project Midterm Review

Nov. 17, 10 AM. Presentation of 5 minutes for individual projects and 10 minutes for group projects: 2 slides on literature survey, 2 slides on progress thus far and 1 slide on future plan of the work.

Logistics

Instructor

Dr. Sriram Ganapathy

sriramg@iisc.ac.in

Office: C 334 (2nd Floor)

Class Times

Mon & Wed

3:45 PM – 5:15 PM

Sun (Tentative)

8:30 AM – 10:00 AM

Where: EE B308 (moved to B-303 EE)

Teaching Assistant

Akshara Soman

aksharas aT iisc doT ac doT in

Lab: C 328 (2nd Floor)

Syllabus

  • •Introduction to real world signals - text, speech, image, video.
  • •Feature extraction and front-end signal processing - information rich representations, robustness to noise and artifacts.
  • •Basics of pattern recognition, Generative modeling - Gaussian and mixture Gaussian models, hidden Markov models.
  • •Discriminative modeling - support vector machines, neural networks and back propagation.
  • •Introduction to deep learning - convolutional and recurrent networks, pre-training and practical considerations in deep learning, understanding deep networks.
  • •Deep generative models - Autoencoders, Boltzmann machines, Adversarial Networks, Variational Learning.
  • •Applications in NLP, computer vision and speech recognition.

Grading Details

15%
Assignments
20%
Midterm Exam
30%
Project
35%
Final Exam

Pre-requisites

Must - Random Process/Probability and Statistics
Must - Linear Algebra/Matrix Theory
Preferred - Basic Digital Signal Processing/Signals and Systems

Textbooks

B1

Pattern Recognition and Machine Learning

C.M. Bishop, 2nd Edition, Springer, 2011.

B2

Neural Networks

C.M. Bishop, Oxford Press, 1995.

B3

Deep Learning

I. Goodfellow, Y. Bengio, A. Courville, MIT Press, 2016.

HTML Version
B4

Fundamentals of Speech Recognition

L. Rabiner, H. Juang, Prentice Hall, 1993.

References

Deep Learning: Methods and Applications

Li Deng, Microsoft Technical Report.

Automatic Speech Recognition - Deep Learning Approach

D. Yu, L. Deng, Springer, 2014.

Machine Learning for Audio, Image and Video Analysis

F. Camastra, Vinciarelli, Springer, 2007.

PDF
Various Published Papers and Online Material
Python Programming Basics
PDF

Slides

06-08-2018
Introduction to real world signals - text, speech, image, video. Learning as a pattern recognition problem. Examples. Roadmap of the course.
08-08-2018
Basics of Natural Language Processing - token, document and corpus. TF-IDF features. Language modeling. Smoothing and back-off. Introduction to audio signal processing. DFT, STFT.
Refs: Information Extraction Book (Ch. 6) - TF-IDF · Stanford Reading Material - Language Modeling
13-08-2018
Revisiting text processing. Perplexity. Short-term Fourier Transform considerations. Time-frequency resolution. Mel-frequency cepstral coefficient (MFCC) features. Image processing - filtering, convolutions. Matrix derivatives.
Refs: Columbia Univ. STFT Tutorial · PRML - Bishop (Appendix, Ch. 12)
20-08-2018
Unsupervised dimensionality reduction using Principal Component Analysis. Maximum variance formulation. Solution using eigenvectors of data covariance matrix. Minimum error formulation. Whitening and standardization. PCA for high dimensional data.
Refs: PRML - Bishop (Ch. 12.1)
27-08-2018
Supervised dimensionality reduction using linear discriminant analysis (LDA). Fisher discriminant. Solution for 2 class LDA. Multi-class LDA. Comparison between PCA and LDA. Introduction to basics of decision theory. Inference and decision problems. Prior, likelihood and posterior. Maximum-a-posteriori decision rule for two class example.
Refs: PRML - Bishop (Ch. 4.1.4, Ch. 1.5)
29-08-2018
Decision theory for regression. MMSE estimation. Multi-variate Gaussian Modeling. Interpretation of Covariance. Diagonal and Full Covariance. Maximum Likelihood estimation of mean and covariance.
Refs: PRML - Bishop (Ch. 1.6)
30-08-2018
Assignment #1. Due on 10-09-2018. Analytical part submitted in class. Coding part submitted via mlsp18.iisc aT gmail doT com.
31-08-2018
Short-comings of single Gaussian modeling. Introduction to mixture Gaussian modeling. Properties and parameters. Expectation Maximization algorithm - auxiliary function, proof of convergence.
Refs: Tutorial on GMMs · Proof of EM algorithm
10-09-2018
Expectation Maximization Algorithm for GMMs. Initialization using K-means. Other Considerations in GMMs. GMM example for unsupervised clustering.
Refs: EM algorithm for GMMs
12-09-2018
Limitations of GMM modeling for sequence data. Markov Chains. Hidden Markov Model (HMM) definition. Three Problems in HMM. Evaluating the likelihood using HMM (Problem 1), Complexity reduction using forward variable and backward variable.
Refs: "Fundamentals of Speech Recognition" - Rabiner and Juang (Ch. 6) · Rabiner Tutorial on HMMs
13-09-2018
Assignment #2. Analytical part submitted in class (26-09-2018). Coding part submitted via mlsp18.iisc aT gmail doT com (28-09-2018).
17-09-2018
Assignment #1 Discussion.
19-09-2018
Inferring the best state alignment - Viterbi algorithm for HMM (Solution to Problem II). Training of HMM using EM algorithm - Baum Welch Algorithm (Solution to Problem III).
Refs: "Fundamentals of Speech Recognition" - Rabiner and Juang (Ch. 6) · Rabiner Tutorial on HMMs
21-09-2018
EM algorithm for HMM with GMM state distribution. Dealing with multiple observation sequences. Implementation issues in HMM. Applications of HMMs - action recognition, face emotion tracking.
Refs: "Fundamentals of Speech Recognition" - Rabiner and Juang (Ch. 6) · Video analysis with HMMs
24-09-2018
Probabilistic PCA. Problem formulation - generative model of the data. EM algorithm for parameter estimation. Application of PPCA.
Refs: PRML - Bishop (Ch. 12.2)
01-10-2018
Regularized linear Regression revisited - dual problem formulation. Gram Matrix. Kernel functions. Examples.
Refs: PRML - Bishop (Ch. 6)
03-10-2018
Mid-term Exam.
10-10-2018
Properties of kernel functions. Rules for constructing kernels. The RBF kernel. Maximum margin classifiers - problem formulation for linearly separable case. Optimization fundamentals - primal and dual problems, strong duality, KKT conditions. Application of KKT conditions to maximum margin classifiers.
Refs: PRML - Bishop (Ch. 7) · Introduction to Convex Optimization - Boyd (Ch. 5)
12-10-2018
Assignment #3. Analytical part submitted in class and coding part submitted via mlsp18.iisc aT gmail doT com (22-10-2018).
12-10-2018
Maximum margin classifiers - overlapping class distribution. Slack variables. Primal and Dual formulation. KKT conditions. Applications of SVM for text classification, cancer detection, MNIST.
Refs: PRML - Bishop (Ch. 7)
15-10-2018
Support vector regression - primal and dual, KKT conditions. Introduction to artificial neural networks - extension of kernel machines. Perceptron model. Learning rule in perceptron. Multi-layer perceptron.
Refs: PRML - Bishop (Ch. 7) · NNPR - Bishop (Ch. 3, 4)
17-10-2018
Forward pass in MLP. Backpropagation algorithm - recursion. Choice of hidden layer activation function.
Refs: NNPR - Bishop (Ch. 4)
22-10-2018
Computational complexity in Gradient Descent. Definition of Jacobian and Hessian matrices. Approximation to Hessian matrix computation. Choice of error function. Mean square and conditional expectation. Conditional expectation for classification with one-hot-encoding. Neural networks estimate posterior probabilities.
Refs: NNPR - Bishop (Ch. 6)
24-10-2018
Assignment #4. Analytical part submitted in class and coding part submitted via mlsp18.iisc aT gmail doT com.
24-10-2018
Cross entropy for two class. Expected cross entropy loss and posterior probability estimation. General condition on error function for outputs to be posterior probability. Weight learning - gradient descent method. Properties of gradient descent using quadratic approximation. Learning rate parameter.
Refs: NNPR - Bishop (Ch. 6, 7)
29-10-2018
Drawbacks of gradient descent. Momentum in learning. Nesterov Accelerated gradient. Second order learning methods. Approximate Hessian. Data preprocessing for Neural networks. Decomposing the error into bias and variance.
Refs: NNPR - Bishop (Ch. 7, 9)
31-10-2018
Improving generalization in deep learning. Regularization - weight decay. Impact of regularization on weight update. Early stopping. Training with noise. Committee of Neural networks. Need for deep architectures.
Refs: NNPR - Bishop (Ch. 9)
02-11-2018
Convolutional Neural Networks. Computation of convolutions. Number of parameters. Advantages over deep neural networks. Pooling and subsampling. Backpropagation in convolution. Insights in deep convolutional networks.
Refs: Deep Learning Book - Goodfellow et al. (Ch. 9)
05-11-2018
Recurrent Neural Networks. Backpropagation in time. Problem of vanishing gradients. Long short term memory networks. Various RNN architectures and applications.
Refs: Deep Learning Book - Goodfellow et al. (Ch. 10)
09-11-2018
Mid-term Exam 2.
12-11-2018
Deep generative modeling - Restricted Boltzmann Machines (RBMs). Conditional independence property. RBM parameter learning. Positive and negative phase of learning. Intuitions behind contrastive divergence algorithm.
Refs: Deep Learning Book - Goodfellow et al. (Ch. 18, 19, 20)
14-11-2018
Assignment #5. Analytical part and coding part submitted via mlsp18.iisc aT gmail doT com (23-11-2018).
14-11-2018
Autoencoders. Denoising AE. Variational autoencoders. Variational lower bound derivation. KL divergence derivation. Data generation with VAEs.
Refs: Kingma's paper · VAE Tutorial
19-11-2018
Generative Adversarial Nets (GANs), Attention Networks. Summary of MLSP course.