Bioinformatics Notes

Machine Learning and Deep Learning for Biological Data

Lecture slides and notes for Machine Learning and Deep Learning for Biological Data in Bioinformatics Notes by Md Ahbab. 13 pages.

Document Info: 13 pages · PDF

Machine Learning and Deep Learning for Biological Data, first page preview

Content Preview

CSE 4893 Introduction to Bioinformatics Module 11 CSE 4893: Introduction to Bioinformatics Module 11: Machine Learning and Deep Learning for Biological Data Abstract These subsidiary notes introduce supervised machine learning and deep learning forbiological data. We begin with what makes measurements of genomes, transcripts and proteins unusual:far more features than samples, batch effects, samples that are related through shared ancestry, anda shortage of trustworthy labels. We then cover representations, classical model families, losses,honest evaluation, and the leakage traps that quietly inflate reported accuracy. The sec- ond halftreats convolutional and attention based sequence models, interpretation methods, re- produciblepractice, and short summaries of landmark papers. Worked exercises, algorithms and a keywordappendix support revision and self study. Contents 1 Learning Objectives 2 2 What Makes Biological Data Different 2 3 Representations 2 4 Model Families 3 5 Objectives and Losses 3 6 Evaluation 4 6.1 Splitting . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 6.2 Metrics . . . . . . . . . . . . . . . . . . . . . . . .

CSE 4893 Introduction to Bioinformatics Module 11 16 Useful Links 12 1 Learning Objectives LO1. Explain why biological data violates the assumptions that standard machine learning tools quietly make, and name the four main causes. LO2. Choose a sensible representation and model family for a given biological data type and sample size, and justify the choice. LO3. Write down and interpret the regularised logistic objective, the softmax, and multi class cross entropy. LO4. Design an evaluation that is honest under class imbalance and homology, including nested cross validation, MCC and precision recall curves. LO5. Describe how convolutional and attention based sequence models work, and interpret them with in silico mutagenesis while avoiding over claiming mechanism. 2 What Makes Biological Data Different The wide data problem. A typical expression study measures tens of thousands of features on afew dozen samples, so the number of features p is much larger than the number of samples n. Withp ≫n there is always a hyperplane that separates the training labels perfectly, even when thelabels are random. Flexible models therefore memorise. The practical consequence is thatregularisation i

CSE 4893 Introduction to Bioinformatics Module 11 Table 1: Common biological data types and their usual numerical representations. Data type Representation Dimensionality Typical model DNA or RNA sequence One hot matrix 4 × L, or k-mer counts 4L, or 4k 1D convolutional network, transformer Protein sequence One hot 20 × L, or language model embedding 20L, or 640 to 2560 per residue Frozen embedding plus linear head Bulk RNA sequencing Log counts per million per gene ∼20 000 Penalised logistic regression Single cell RNA sequencing Sparse counts, then principal components 20 000 down to 50 Gradient boosting, autoencoder Genetic variants Genotype dosage matrix in {0, 1, 2} 105 to 107 Penalised linear models Mass spectra, metabolites Binned or aligned intensity vector 103 to 104 Random forest, support vector machine Protein structure Contact map or residue graph L × L Graph neural network Table 2: Five model families you should be able to choose between and defend. Model Works well when Fails when Interpretability Penalised logistic regression p ≫n, roughly additive effects, small samples Strong interactions or non linear thresholds matter High, signed coefficients Support vector machin