Machine Learning for Signal Processing

E9 205 • Fall 2016

Announcements

Final Exam

Dec 7, 1:30 PM – 4:30 PM, MMCR. Open book, open notes. No laptops/cellphones allowed.

Take Home Exam 2

Take home exam-2 has been posted.

Project Evaluation

Dec 19, 9:30 AM – 1:00 PM, MMCR.

  • Single person projects (max 10 min presentation, max 5 slides).
  • Multi person projects - both individuals presenting (max 15 min presentation, max 8 slides).
  • Project components - implementation of baseline paper, comparing results with baseline paper, novel directions improving the baseline.
  • Project report (max single column 5 pages) - due Dec 17.
  • Project slides emailed by Dec 18.
  • Mark distribution (Total Marks 30) - Mid-term Evaluation (7), Final presentation (5), Report (5), Baseline implementation (7), Novelty (6).

Logistics

Instructor

Sriram Ganapathy

sriram aT ee doT iisc doT ernet doT in

Office: C 334 (2nd Floor)

Class Times

Mon & Wed

3:30 PM – 5:00 PM

Where: EE C241 (MMCR 1st Floor)

Teaching Assistant

Achuth Rao

achuthraomv aT gmail doT com

Lab: C 326 (2nd Floor)

TA Hours: Thu 3–5 PM

Syllabus

  • •Introduction to real world signals - text, speech, image, video.
  • •Feature extraction and front-end signal processing - information rich representations, robustness to noise and artifacts, signal enhancement, bio inspired feature extraction.
  • •Basics of pattern recognition, generative modeling - Gaussian and mixture Gaussian models, hidden Markov models, factor analysis and latent variable models.
  • •Discriminative modeling - support vector machines, neural networks and back propagation.
  • •Introduction to deep learning - convolutional and recurrent networks, pre-training and practical considerations in deep learning, understanding deep networks.
  • •Clustering methods and decision trees. Feature selection methods.
  • •Applications in computer vision and speech recognition.

Grading Details

15%
Assignments
20%
Midterm Exam
30%
Project
35%
Final Exam

Pre-requisites

Random Process/Probability and Statistics
Linear Algebra/Matrix Theory
Basic Digital Signal Processing/Signals and Systems

Textbooks

B1

Pattern Recognition and Machine Learning

C.M. Bishop, 2nd Edition, Springer, 2011.

B2

Deep Learning

I. Goodfellow, Y. Bengio, A. Courville, MIT Press, 2016.

HTML Version
B3

Digital Image Processing

R.C. Gonzalez, R.E. Woods, 3rd Edition, Prentice Hall, 2008.

B4

Fundamentals of Speech Recognition

L. Rabiner and H. Juang, Prentice Hall, 1993.

References

Deep Learning: Methods and Applications

Li Deng, Microsoft Technical Report.

Automatic Speech Recognition - Deep Learning Approach

D. Yu, L. Deng, Springer, 2014.

Machine Learning for Audio, Image and Video Analysis

F. Camastra, Vinciarelli, Springer, 2007.

PDF

Slides

03-08-2016
Introduction to real world signals - text, speech, image, video. Learning as a pattern recognition problem. Examples. Roadmap of the course.
08-08-2016
Types of learning methods, feature extraction for speech and audio, short-term Fourier transform, narrow band and wideband spectrogram, time frequency resolution.
Refs: Dan Ellis - Tutorial · Ricardo - Tutorial
10-08-2016
Uncorrelated noise in speech/audio, non-negative matrix factorization (NMF), problem definition, cost function and constraints, auxiliary function, proof of convergence, parameter update rule. Application to audio source separation and speech denoising.
Refs: Bhiksha Raj - Tutorial · Lee - Paper
17-08-2016
Linear Prediction - orthogonality of prediction error with past samples, optimal linear predictor, Yule-Walker equations, energy of prediction error, stability of prediction filter, autoregressive process, linear prediction for AR process.
Ref: Theory of LP - Vaidyanathan (Chap. 2, 5.3, A, B)
22-08-2016
Normal equations for autoregressive process. Power spectral density. Autoregressive modeling of PSD. Applications of linear prediction.
Assignment #1. Non-negative Matrix Factorization, Linear Prediction, applications for face images and noisy speech. Due 02-09-2016 (noon).
24-08-2016
Matrix derivative rules. Dimensionality reduction I - Principal component analysis (PCA), maximum variance formulation, minimum error formulation. Whitening and standardization, PCA for high dimensional data. Linear discriminant analysis (LDA), Fisher discriminant for two classes.
Ref: PRML - Bishop
29-08-2016
LDA for multiple classes, LDA formulation in lower dimensional subspace. Applications of PCA. Distinction between PCA and LDA. Introduction to feature extraction from image data - wavelet transform, mother wavelet, scaling and shifting, continuous and dyadic wavelet transform.
Ref: Introduction to Wavelets and Wavelet Transforms - Burrus et al.
31-08-2016
Dyadic wavelet transform. Scaling and wavelet function. Approximation and detail. Wavelet decomposition. Application to 1-D signals.
Ref: Selected Pages - Burrus et al. (Chap. 2)
05-09-2016
Interpreting wavelet approximation and detail coefficients. Filter bank approach to wavelets. Extension to 2-D wavelet transform, application to images.
Refs: Tutorial on 2-D Wavelets · Image Denoising
07-09-2016
Decision theory - inference and decision rule, mis-classification error, maximum posterior decision rule, expected loss, minimum mean square error decision rule for regression. Three approaches to inference and decision - generative modeling, discriminative modeling and discriminant functions.
Ref: PRML - Bishop (Sec. 1.5)
12-09-2016
Introduction to generative modeling. Gaussian distribution. Parameter estimation using maximum likelihood (MLE). Sample mean and sample covariance. Limitations of Gaussian modeling. Gaussian mixture model (GMM) density function.
14-09-2016
MLE for GMM - Expectation Maximization (EM) algorithm. Proof of EM algorithm. Convergence properties. EM algorithm for GMM parameter estimation. Choice of hidden variable. Application of GMMs for unsupervised data clustering.
Refs: Tutorial GMMs · Proof of EM Algorithm · EM Algorithm for GMMs
19-09-2016
Markov chain - sequence modeling with hidden Markov modeling (HMM). Definition of HMM parameters. Three problems in HMM - (i) evaluation (ii) inference and (iii) training. Direct computation of likelihood. Forward and backward variable recursion.
Refs: Rabiner, Juang, "Fundamentals of Speech Recognition" (Chap. 6) · SP Magazine Article - Rabiner
21-09-2016
Solution to problem (ii) in HMM - Viterbi algorithm. HMM parameter estimation with EM algorithm. Estimation of Q function and iterative model update.
Refs: Tutorial HMMs · Rabiner, Juang, "Fundamentals of Speech Recognition" (Chap. 6)
Assignment #2 (Part A). PCA/LDA, ML, Gaussian and GMM, HMM. Due 03-10-2016 (class).
26-09-2016
First Mid-term Exam.
28-09-2016
Discussion on first mid-term exam. Topics for mini-projects.
03-10-2016
Hidden Markov Models with GMM observation densities. Application of EM algorithm for GMM-HMMs. Parameter estimation. Application of HMMs in video analysis. Dimensionality reduction continued - latent variable models.
Refs: GMM-HMM - "Fundamentals of Speech Recognition", Rabiner (Chap. 6) · Slides from N. Ramanathan - Video Analysis with HMMs
05-10-2016
Probabilistic PCA (PPCA) - generative model description. Log-likelihood computation, parameter estimation using direct optimization. EM algorithm for PPCA. Extension to factor analysis. Summary of generative modeling. Introduction to discriminative modeling - non-linear regression with kernels.
Ref: PRML - Bishop (Sec. 12.2, 3.1) · Paper: "PPCA", Tipping et al.
12-10-2016
Recap of generative versus discriminative modeling. Non-linear regression with regularization. Dual problem definition and solution with kernels. Properties of kernel functions. Constructing kernels from basic blocks. Sparse kernel machines.
Ref: PRML - Bishop (Sec. 3.3, 6)
17-10-2016
Classifiers with kernels. Definition of margin. Maximum margin classifiers. Introduction to convex optimization with constraints. Primal and dual problems. Weak and strong duality. Karush-Kuhn-Tucker (KKT) conditions for strong duality. Solving the dual problem for maximum margin classifiers. Definition of support vectors.
Ref: PRML - Bishop (Chap. 7.1) · Book (Chap. 5): "Convex Optimization", Boyd and Vandenberghe
Assignment #2 (Part B). Implementing PCA/LDA, GMM and HMM. Due 28-10-2016 (noon).
19-10-2016
Maximum margin classifiers for overlapping class distributions, concept of slack variables. Lagrangian and dual form. KKT conditions for solving the optimal parameters. Sequential minimal optimization algorithm - analytic solution to two variable constrained optimization problem, heuristics for choosing the two variables. Estimating the bias parameter in SVM.
Ref: PRML - Bishop (Chap. 7.1.1) · SMO Paper - J. Platt et al.
24-10-2016
Summary of support vector machines - problem definition, primal and dual formulations, kernel space transformation, solutions and implications, applications of SVMs in cancer diagnosis and text categorization. Support vector regression - slack variables and dual formulation.
Ref: PRML - Bishop (Chap. 7.1.4) · NYU Bio Medicine - Tutorial
31-10-2016
Introduction to neural networks. Illustration with XOR problem - need for hidden layer(s) with non-linear activations. Optimization methods for neural networks. First order Taylor series - gradient descent method. Curvature and second derivatives. Jacobian and Hessian matrices. Newton's method. Stochastic gradient descent.
Ref: DLB (Deep Learning Book) - Goodfellow, Bengio (Chap. 6, 4.3)
02-11-2016
Neural networks estimate posterior probabilities. Architecture considerations - cost function (mean square error, cross entropy), output units (linear, sigmoidal or softmax), hidden unit activations (ReLU and variants, tanh or sigmoidal).
Ref: DLB (Deep Learning Book) - Goodfellow, Bengio (Chap. 6)
07-11-2016
Universal approximation properties of NNs. Need for multiple hidden layers. Depth versus width. Mechanism of representation learning in deep networks. Parameter learning in deep networks - back propagation. Equivalence in learning DNNs with linear output activation and MSE versus softmax activations with cross entropy error.
Refs: DLB (Deep Learning Book) - Goodfellow, Bengio (Chap. 6) · ASR - DL Approach, D. Yu, Li Deng (Chap. 4)
09-11-2016
Summary of NN learning and architecture. Pseudo code for back propagation. Other considerations - data preprocessing, model initialization. Underfit versus overfit. Improving generalization with regularization. L2 regularization. Quadratic approximation.
Ref: DLB (Deep Learning Book) - Goodfellow, Bengio (Chap. 6, 7)
14-11-2016
L1 regularization. Multi-task learning. Early stopping. Equivalence between L2 regularization and early stopping. Bagging and ensemble averaging. Dropout.
Ref: DLB (Deep Learning Book) - Goodfellow, Bengio (Chap. 7)
16-11-2016
Convolutional neural networks. Filtering and hierarchical sparsity. Pooling and striding. Deep convolutional networks. Discussion on second mid-term exam.
Ref: DLB (Deep Learning Book) - Goodfellow, Bengio (Chap. 9)
21-11-2016
Deep generative models - Restricted Boltzmann Machine, model definition, conditional independence. Relationship with sigmoidal activation. Parameter learning in RBM - positive and negative phase, approximation with sampling methods, contrastive divergence algorithm. Deep Belief Networks (DBNs).
Ref: DLB (Deep Learning Book) - Goodfellow, Bengio (Chap. 18, 20)
23-11-2016
Gaussian Restricted Boltzmann Machine (GRBM). Relationship with GMMs. Summary of deep learning methods.
Ref: ASR - Deep Learning Approach (Yu and Deng) - Chap. 5
29-11-2016
Take Home Practice Exam.