Novel statistical modeling and selection methodologies for high dimensional genomic data

Novel statistical modeling and selection methodologies for high dimensional genomic data

by David Andrew Engler

About
This work presents new contributions to the analysis of microarray data. Chapter 1 and Chapter 2 address scientific questions of interest specific to array-based comparative genomic hybridizations (aCGH). In Chapter 1, an approach for the identification of DNA copy number gains and losses is presented. The approach borrows strength across chromosomes and across hybridizations. Additionally, the employed model allows for intertumoral variation, as well as intratumoral clonal variation. The method produces quantitative assessments of the likelihood of genetic alterations at each clone, along with a graphical display for simple visual interpretation. In Chapter 2, a novel method for the assessment of differences in aCGH-based genetic instability across cancer subtypes is presented. Instability phenotypes are composed of a variety of copy number alteration features including height or magnitude of copy number alteration, frequency of transition between copy number states such as gain and loss, and total number of altered clones or probes. That is, instability phenotype is multivariate in nature. Current methods of instability phenotype assessment, however, are limited to univariate measures and are therefore limited in both accuracy and interpretability. The proposed method is based on the pseudolikelihood model of Chapter 1. Through use of a pseudolikelihood ratio test (PLRT), more accurate assessment of instability phenotype differences between cancer subtypes is possible. In Chapter, 3 a method is presented for the identification of genes associated with survival. The method is broadly applicable to aCGH data as well as to data obtained from other microarray platforms. Adaptation of the elastic net penalized variable selection approach is presented under the Cox proportional hazards model and under an accelerated failure time (AFT) model. The proposed methods provide computationally efficient approaches with predictive performance superior to existing methods when identification of highly correlated variables is of interest and in settings where a high degree of censoring is present.

Discuss Novel statistical modeling and selection methodologies for high dimensional genomic data with other readers

Join or start a book club for Novel statistical modeling and selection methodologies for high dimensional genomic data on Readfeed. Live chat, shared reading progress, and AI discussion questions — free to get started.

Frequently asked questions

How do I join a book club for Novel statistical modeling and selection methodologies for high dimensional genomic data?

Sign up free on Readfeed, then browse public clubs or start your own club with Novel statistical modeling and selection methodologies for high dimensional genomic data as the current read. Invite friends with a share link and discuss together with live chat and AI discussion questions.

Can I discuss Novel statistical modeling and selection methodologies for high dimensional genomic data with other readers online?

Yes. Readfeed book clubs let you chat live, share progress, and join discussions about Novel statistical modeling and selection methodologies for high dimensional genomic data with readers worldwide — whether your club is virtual, in-person, or hybrid.

Is Readfeed free?

Yes. Creating an account and joining book clubs is free. Sign up to find readers who love the same books and start discussing today.