Machine Learning for Signal Processing

E9 205 • Fall 2017

Announcements

Final Exam

10-12-2017 (2:00 PM – 5:00 PM), B303 (Classroom). Open book, open notes. No laptops/cellphones allowed.

Project Evaluation

15-12-2017 (9:00 AM). Maximum 8 slides (single person) or 12 slides (2 persons) per project. Maximum 3-4 pages report (submit report and slides by noon Dec. 14 through mail). Evaluation criteria: focus on problem definition and motivation, implementing the baseline and your contribution.

Feedback Form

Fifth Assignment

Due 24-11-2017.

HW5

Logistics

Instructor

Sriram Ganapathy

sriram aT ee doT iisc doT ernet doT in

Office: C 334 (2nd Floor)

Class Times

Mon & Wed

3:30 PM – 5:00 PM

Where: EE B303 (second class onwards)

Teaching Assistant

Aravind Illa

aravindece77 aT gmail doT com

Lab: C 326 (2nd Floor)

Syllabus

  • •Introduction to real world signals - text, speech, image, video.
  • •Feature extraction and front-end signal processing - information rich representations, robustness to noise and artifacts, signal enhancement, bio inspired feature extraction.
  • •Basics of pattern recognition, Generative modeling - Gaussian and mixture Gaussian models, hidden Markov models, factor analysis.
  • •Discriminative modeling - support vector machines, neural networks and back propagation.
  • •Introduction to deep learning - convolutional and recurrent networks, pre-training and practical considerations in deep learning, understanding deep networks.
  • •Deep generative models - Autoencoders, Boltzmann machines, Adversarial Networks.
  • •Applications in computer vision and speech recognition.

Grading Details

15%
Assignments
20%
Midterm Exam
30%
Project
35%
Final Exam

Pre-requisites

Random Process/Probability and Statistics
Linear Algebra/Matrix Theory
Basic Digital Signal Processing/Signals and Systems

Textbooks

B1

Pattern Recognition and Machine Learning

C.M. Bishop, 2nd Edition, Springer, 2011.

B2

Neural Networks

C.M. Bishop, Oxford Press, 1995.

B3

Deep Learning

I. Goodfellow, Y. Bengio, A. Courville, MIT Press, 2016.

HTML Version
B4

Digital Image Processing

R. C. Gonzalez, R. E. Woods, 3rd Edition, Prentice Hall, 2008.

B5

Fundamentals of Speech Recognition

L. Rabiner and H. Juang, Prentice Hall, 1993.

References

Deep Learning: Methods and Applications

Li Deng, Microsoft Technical Report.

Automatic Speech Recognition - Deep Learning Approach

D. Yu, L. Deng, Springer, 2014.

Machine Learning for Audio, Image and Video Analysis

F. Camastra, Vinciarelli, Springer, 2007.

PDF

Slides

14-08-2017
Introduction to real world signals - text, speech, image, video. Learning as a pattern recognition problem. Examples. Roadmap of the course.
16-08-2017
Feature Extraction - Goals and challenges. Introduction to text processing. Bag of words model. Term Frequency-Inverse document frequency. N-gram modeling. Feature Extraction in Audio and Speech - Spectrogram.
21-08-2017
Mel-frequency cepstral Coefficients (MFCC), Linear Prediction - orthogonality of prediction error with past samples, optimal linear predictor, stability of prediction filter, Autoregressive process, linear prediction for AR process.
23-08-2017
Basics for Digital Image Processing – Filtering, Smoothing, Edge Detection, Scale Invariant Feature Transform (SIFT).
28-08-2017
Matrix and vector derivatives - definition and properties. Dimensionality reduction - Preserving maximum data variance - principal component analysis (PCA). Minimum error formulation of PCA. Residual error in PCA. Example of PCA application for hand-written digit images.
Refs: PRML - Bishop (Appendix, Chapter 12)
30-08-2017
PCA for high dimensional data. Whitening and KL transform. Limitations of PCA. Class dependent dimensionality reduction using linear discriminant analysis (LDA). Fisher discriminant for 2 class case using within-class and between class matrices. Solution of LDA. Multi-class LDA, PCA versus LDA example.
Refs: PRML - Bishop (Chapter 4.1.4)
01-09-2017
Basics of Python Programming. Installing python, simple commands and functions. Loading speech and image data. Vectorizing, mean computation and spectrogram.
01-09-2017
Assignment #1. Due on 11-09-2017. Analytical part submitted in class. Coding part submitted via e9205mlsp2017 aT gmail doT com.
04-09-2017
Decision theory basics. Minimum classification error rule. MAP and ML based approaches. 3 approaches to ML. Generative versus discriminative modeling. Introduction to generative modeling. Multi-variate Gaussian Distribution.
Refs: PRML - Bishop (Chapter 1.5)
06-09-2017
MLE for multi-variate Gaussian. Sample mean and variance. Limitations of Gaussian modeling. Need for mixture modeling. Probability density of Gaussian Mixture Model (GMM).
11-09-2017
MLE for GMM - Expectation Maximization (EM) algorithm. Proof of EM algorithm. Convergence properties. EM algorithm for GMM parameter estimation. Choice of hidden variable.
Refs: Tutorial on GMMs · Proof of EM algorithm · EM algorithm for GMMs
13-09-2017
Summary of GMM modeling. Application of GMM for unsupervised clustering.
18-09-2017
Limitations of GMM modeling for sequence data. Markov Chains. Hidden Markov Model (HMM) definition. Three Problems in HMM.
Refs: "Fundamentals of Speech Recognition", Rabiner and Juang (Chapter 6)
20-09-2017
Evaluating the likelihood using HMM (Problem 1), Complexity reduction using forward variable and backward variable. Finding the best state sequence (Problem 2) - instantaneous probability based, Viterbi algorithm for state sequence segmentation.
Refs: "Fundamentals of Speech Recognition", Rabiner and Juang (Chapter 6)
23-09-2017
Re-estimating the HMM parameters - EM algorithm for HMM (Problem 3). Q function definition and solution. Intuitions about HMM training.
Refs: "Fundamentals of Speech Recognition", Rabiner and Juang (Chapter 6) · EM algorithm for HMMs
25-09-2017
Non-negative matrix factorization (NMF), problem definition, cost function and constraints, auxiliary function, proof of convergence, parameter update rule. Application to audio source separation and speech denoising.
Refs: Bhiksha Raj Tutorial · Lee Paper on NMFs
04-10-2017
First Mid-term Exam.
09-10-2017
Application of NMF. Audio separation into individual instruments, speech denoising with known and unknown sources. Linear models for regression - problem definition. Least squares regression. Maximum likelihood and least squares regression.
Refs: PRML - Bishop (Chapter 2)
11-10-2017
Overfitting and Underfitting. Regularized least squares. Linear Models for Classification. Least squares for classification. Sigmoid function and one-of-K encoding. Problems with least squares classification.
Refs: PRML - Bishop (Chapter 3)
16-10-2017
Logistic regression - two class problem. Sigmoid function and posterior probability. Logistic regression - K class problem. Softmax function and cross entropy error function and Maximum likelihood estimation. Linear regression revisited - dual formulation.
Refs: PRML - Bishop (Chapter 4, 6)
21-10-2017
Design matrix, kernel function and Gram matrix. Necessary and sufficient condition for kernel functions (Mercer's theorem). Examples of kernel functions.
Refs: PRML - Bishop (Chapter 6)
23-10-2017
Margin of linear classifier. Maximum margin classifier formulation. Constraints involved in optimization. Introduction to support vector machines.
Refs: PRML - Bishop (Chapter 7)
25-10-2017
Introduction to constrained optimization. Primal and dual problems. Weak and strong duality. Necessary and sufficient conditions for strong duality for convex problems with convex conditions. KKT conditions.
Refs: Introduction to Convex Optimization - Boyd (Chapter 5)
27-10-2017
Application of convex optimization to SVMs. KKT conditions and solution to problem. Definition of support vectors. Support vector machine for overlapping classes. Trade off in regularization and training loss.
Refs: PRML - Bishop (Chapter 7)
30-10-2017
SVM application for classification. Support vector regression. Formulation and KKT conditions. Introduction to neural networks. Parameter learning using gradient descent (scalar case).
Refs: PRML - Bishop (Chapter 7)
03-11-2017
Gradient descent vector case. Types of activation functions. XOR problem with NNs. Need for deep architecture neural networks.
Refs: Deep Learning - Goodfellow et al. (Chapter 6)
04-11-2017
Learning in Neural networks. First order methods - Method of steepest descent. Curvature and Hessians. Second order method - Newton method. Discussion on complexity of learning algorithms.
Refs: Deep Learning - Goodfellow et al. (Chapter 4), Neural Networks - Bishop (Chapter 4, 7)
06-11-2017
Back propagation algorithm for learning in deep networks. Linear neuron with MSE algorithm. Disadvantages and limitations of gradient descent algorithm.
Refs: Neural Networks - Bishop (Chapter 6)
08-11-2017
Second Mid-term Exam.
10-11-2017
Types of non-linearities used. Cost function for regression and classification. Output activation function used in regression and classification. Equivalence between regression with MSE and classification with CE using softmax output activations.
Refs: Neural Networks - Bishop (Chapter 6)
13-11-2017
Learning and Generalization issues in Neural networks. Decomposing the MSE into bias and variance. Discussion on bias variance tradeoff. Improving learning with regularization.
Refs: Neural Networks - Bishop (Chapter 9)
14-11-2017
Assignment #5. Due on 24-11-2017. Analytical part submitted in class. Coding part submitted via e9205mlsp2017 aT gmail doT com.
15-11-2017
L2 weight regularization, early stopping and training with added noise in the input data. Committees of neural networks. System combination methods and optimization.
Refs: Neural Networks - Bishop (Chapter 9)
17-11-2017
Improving the speed of convergence of gradient descent with momentum. Convolutional neural networks. Kernels, pooling and sub-sampling. Comparison of CNNs and DNNs. Weight sharing and parameter learning.
Refs: Deep Learning - Goodfellow et al. (Chapter 9)
18-11-2017
Understanding the learning in deep layers of CNNs. Recurrent networks. Backpropagation in time for RNN parameter learning. Various RNN architectures - teacher forcing, sequence-to-vector and bi-directional RNNs.
Refs: Deep Learning - Goodfellow et al. (Chapter 10)
20-11-2017
Long short term memory networks. Deep unsupervised learning - Restricted Boltzmann Machines (RBMs). Conditional independence in RBMs. Learning in RBM with maximum likelihood. Positive and negative partition function. Gibbs sampling and contrastive divergence approximations.
Refs: Deep Learning - Goodfellow et al. (Chapter 18, 20)