Statistical analysis and modeling of mass spectrometry-based metabolomics data Article

Xi, B, Gu, H, Baniasadi, H et al. (2014). Statistical analysis and modeling of mass spectrometry-based metabolomics data . 1198 333-353. 10.1007/978-1-4939-1258-2_22

cited authors

  • Xi, B; Gu, H; Baniasadi, H; Raftery, D

authors

abstract

  • Multivariate statistical techniques are used extensively in metabolomics studies, ranging from biomarker selection to model building and validation. Two model independent variable selection techniques, principal component analysis and two sample t-tests are discussed in this chapter, as well as classification and regression models and model related variable selection techniques, including partial least squares, logistic regression, support vector machine, and random forest. Model evaluation and validation methods, such as leave-one-out cross-validation, Monte Carlo cross-validation, and receiver operating characteristic analysis, are introduced with an emphasis to avoid over-fitting the data. The advantages and the limitations of the statistical techniques are also discussed in this chapter.

publication date

  • January 1, 2014

Digital Object Identifier (DOI)

start page

  • 333

end page

  • 353

volume

  • 1198