Bioinformatics Notes

Functional Enrichment, Pathways and Network Biology

Lecture slides and notes for Functional Enrichment, Pathways and Network Biology in Bioinformatics Notes by Md Ahbab. 12 pages.

Document Info: 12 pages · PDF

Functional Enrichment, Pathways and Network Biology, first page preview

Content Preview

CSE 4893 Introduction to Bioinformatics Module 8 CSE 4893: Introduction to Bioinformatics Module 8: Functional Enrichment, Pathways and Network Biology Abstract These subsidiary notes cover the interpretation stage of a genomics experiment, where a filtered or ranked gene list has to become a biological story. We define gene sets and survey the main annotation sources, derive over representation analysis from the hypergeometric distribution, and build gene set enrichment analysis from a weighted running sum with permutation based error control. We then move from single gene sets to whole networks, covering degree, clustering, modularity and centrality, the common network types and their biases, community detection with Louvain and Leiden, and weighted co-expression networks. Worked examples, algorithms, reporting guidance and a keyword appendix support laboratory practice and revision. Contents 1 Learning Objectives 1 2 Gene Sets and Their Sources 2 3 Over Representation Analysis 2 4 Gene Set Enrichment Analysis 3 5 Graph Foundations 4 6 Network Types 5 7 Modules and Communities 5 8 Weighted Co-expression Networks 6 9 PLACEHOLDER: Additional Formulas 6 10 Algorithms 6 11 Reporting

CSE 4893 Introduction to Bioinformatics Module 8 2. Carry out and interpret an over representation analysis using the hypergeometric test, in- cluding the choice of background universe and multiple testing correction. 3. Describe the gene set enrichment analysis running sum, the enrichment score, the normalised enrichment score and the permutation based false discovery rate. 4. Compute and interpret basic network measures such as degree, clustering coefficient, between- ness centrality and modularity, and recognise a scale free degree distribution. 5. Build a weighted co-expression network, detect modules, relate module eigengenes to traits, and report the analysis honestly with its biases. 2 Gene Sets and Their Sources A gene set is simply a named list of genes that share something biologically meaningful, such as a molecular function, a pathway membership or a shared response to a treatment. Enrichment analysis asks whether your experimental gene list overlaps such a set more than chance would allow. Table 1 compares the sources you are most likely to use, with release information current at the time of writing. Table 1: Major gene set and pathway sources, with structure, cadence

CSE 4893 Introduction to Bioinformatics Module 8 Hypergeometric probability and tail p value P(X = k) = K k N−K n−k  N n  , P(X ≥k) = min(K,n) X i=k K i N−K n−i  N n  N size of the background universe, that is every gene that could have been measured and annotated. K number of genes in the universe that belong to the gene set under test. n size of your selected list, after intersecting it with the universe. k number of genes that are both in your list and in the gene set, the observed overlap. 3.1 A worked example Suppose an RNA sequencing study measures N = 20,000 genes, of which K = 200 belong to the pathway “cell cycle checkpoint”. You call n = 300 genes differentially expressed and find k = 15 of them in that pathway. The expected overlap is nK/N = 3, so the fold enrichment is 15/3 = 5, and the tail probability is about 6.7×10−7. Because you test hundreds or thousands of sets at once, raw Table 2: The two by two table behind the worked over representation example. In selected list Not selected Total In gene set k = 15 185 K = 200 Not in gene set 285 19,515 19,800 Total n = 300 19,700 N = 20,000 p values must be adjusted. The Benjamini and Hochberg procedure controls