Surveying the manifold divergence of an entire protein class for statistical clues to underlying biochemical mechanisms.

Abstract:

:Certain residues have no known function yet are co-conserved across distantly related protein families and diverse organisms, suggesting that they perform critical roles associated with as-yet-unidentified molecular properties and mechanisms. This raises the question of how to obtain additional clues regarding these mysterious biochemical phenomena with a view to formulating experimentally testable hypotheses. One approach is to access the implicit biochemical information encoded within the vast amount of genomic sequence data now becoming available. Here, a new Gibbs sampling strategy is formulated and implemented that can partition hundreds of thousands of sequences within a major protein class into multiple, functionally-divergent categories based on those pattern residues that best discriminate between categories. The sampler precisely defines the partition and pattern for each category by explicitly modeling unrelated, non-functional and related-yet-divergent proteins that would otherwise obscure the analysis. To aid biological interpretation, auxiliary routines can characterize pattern residues within available crystal structures and identify those structures most likely to shed light on the roles of pattern residues. This approach can be used to define and annotate automatically subgroup-specific conserved domain profiles based on statistically-rigorous empirical criteria rather than on the subjective and labor-intensive process of manual curation. Incorporating such profiles into domain database search sites (such as the NCBI BLAST site) will provide biologists with previously inaccessible molecular information useful for hypothesis generation and experimental design. Analyses of P-loop GTPases and of AAA+ ATPases illustrate the sampler's ability to obtain such information.

authors

Neuwald AF

doi

10.2202/1544-6115.1666

subject

Has Abstract

pub_date

2011-01-01 00:00:00

pages

Article 36

eissn

2194-6302

issn

1544-6115

journal_volume

10

pub_type

杂志文章
  • Transmission disequilibrium test power and sample size in the presence of locus heterogeneity.

    abstract::Locus heterogeneity is one of the most important issues in gene mapping and can cause significant reductions in statistical power for gene mapping, yet no research to date has provided power and sample size calculations for family-based association methods in the presence of locus heterogeneity. The purpose of this re...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1501

    authors: Chen C,Yang G,Buyske S,Matise T,Finch SJ,Gordon D

    更新日期:2009-01-01 00:00:00

  • Semi-parametric differential expression analysis via partial mixture estimation.

    abstract::We develop an approach for microarray differential expression analysis, i.e. identifying genes whose expression levels differ between two or more groups. Current approaches to inference rely either on full parametric assumptions or on permutation-based techniques for sampling under the null distribution. In some situa...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1333

    authors: Rossell D,Guerra R,Scott C

    更新日期:2008-01-01 00:00:00

  • M-quantile regression analysis of temporal gene expression data.

    abstract::In this paper, we explore the use of M-quantile regression and M-quantile coefficients to detect statistical differences between temporal curves that belong to different experimental conditions. In particular, we consider the application of temporal gene expression data. Here, the aim is to detect genes whose temporal...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1452

    authors: Vinciotti V,Yu K

    更新日期:2009-01-01 00:00:00

  • Fully Bayesian mixture model for differential gene expression: simulations and model checks.

    abstract::We present a Bayesian hierarchical model for detecting differentially expressed genes using a mixture prior on the parameters representing differential effects. We formulate an easily interpretable 3-component mixture to classify genes as over-expressed, under-expressed and non-differentially expressed, and model gene...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1314

    authors: Lewin A,Bochkina N,Richardson S

    更新日期:2007-01-01 00:00:00

  • On the operational characteristics of the Benjamini and Hochberg False Discovery Rate procedure.

    abstract::Multiple testing procedures are commonly used in gene expression studies for the detection of differential expression, where typically thousands of genes are measured over at least two experimental conditions. Given the need for powerful testing procedures, and the attendant danger of false positives in multiple testi...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1302

    authors: Green GH,Diggle PJ

    更新日期:2007-01-01 00:00:00

  • MLML2R: an R package for maximum likelihood estimation of DNA methylation and hydroxymethylation proportions.

    abstract::Accurately measuring epigenetic marks such as 5-methylcytosine (5-mC) and 5-hydroxymethylcytosine (5-hmC) at the single-nucleotide level, requires combining data from DNA processing methods including traditional (BS), oxidative (oxBS) or Tet-Assisted (TAB) bisulfite conversion. We introduce the R package MLML2R, which...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2018-0031

    authors: Kiihl SF,Martinez-Garrido MJ,Domingo-Relloso A,Bermudez J,Tellez-Plaza M

    更新日期:2019-01-17 00:00:00

  • A Bayesian approach to estimation and testing in time-course microarray experiments.

    abstract::The objective of the present paper is to develop a truly functional Bayesian method specifically designed for time series microarray data. The method allows one to identify differentially expressed genes in a time-course microarray experiment, to rank them and to estimate their expression profiles. Each gene expressio...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1299

    authors: Angelini C,De Canditiis D,Mutarelli M,Pensky M

    更新日期:2007-01-01 00:00:00

  • Empirical bayes microarray ANOVA and grouping cell lines by equal expression levels.

    abstract::In the exploding field of gene expression techniques such as DNA microarrays, there are still few general probabilistic methods for analysis of variance. Linear models and ANOVA are heavily used tools in many other disciplines of scientific research. The usual F-statistic is unsatisfactory for microarray data, which e...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1125

    authors: Lönnstedt I,Rimini R,Nilsson P

    更新日期:2005-01-01 00:00:00

  • Accommodating uncertainty in a tree set for function estimation.

    abstract::Multiple branching trees have been used to model the acquisition of HIV drug resistance mutations, and several different algorithms have been developed to construct the tree set that best describes the data. These algorithms have mainly focused on the structure of the tree set. The focal point of this paper is estimat...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1324

    authors: Healy BC,DeGruttola VG,Hu C

    更新日期:2008-01-01 00:00:00

  • Discrete Wavelet Packet Transform Based Discriminant Analysis for Whole Genome Sequences.

    abstract::In recent years, alignment-free methods have been widely applied in comparing genome sequences, as these methods compute efficiently and provide desirable phylogenetic analysis results. These methods have been successfully combined with hierarchical clustering methods for finding phylogenetic trees. However, it may no...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2018-0045

    authors: Huang HH,Girimurugan SB

    更新日期:2019-02-15 00:00:00

  • Variance and covariance heterogeneity analysis for detection of metabolites associated with cadmium exposure.

    abstract::In this study, we propose a novel statistical framework for detecting progressive changes in molecular traits as response to a pathogenic stimulus. In particular, we propose to employ Bayesian hierarchical models to analyse changes in mean level, variance and correlation of metabolic traits in relation to covariates. ...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2013-0041

    authors: Salamanca BV,Ebbels TM,Iorio MD

    更新日期:2014-04-01 00:00:00

  • Combining dependent p-values by gamma distributions.

    abstract::Combining correlated p-values from multiple hypothesis testing is a most frequently used method for integrating information in genetic and genomic data analysis. However, most existing methods for combining independent p-values from individual component problems into a single unified p-value are unsuitable for the cor...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2019-0057

    authors: Chien LC

    更新日期:2020-11-06 00:00:00

  • Reproducibility of biomarker identifications from mass spectrometry proteomic data in cancer studies.

    abstract::Reproducibility of disease signatures and clinical biomarkers in multi-omics disease analysis has been a key challenge due to a multitude of factors. The heterogeneity of the limited sample, various biological factors such as environmental confounders, and the inherent experimental and technical noises, compounded wit...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2018-0039

    authors: Liang Y,Kelemen A,Kelemen A

    更新日期:2019-05-11 00:00:00

  • Approximate maximum likelihood estimation for population genetic inference.

    abstract::In many population genetic problems, parameter estimation is obstructed by an intractable likelihood function. Therefore, approximate estimation methods have been developed, and with growing computational power, sampling-based methods became popular. However, these methods such as Approximate Bayesian Computation (ABC...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2017-0016

    authors: Bertl J,Ewing G,Kosiol C,Futschik A

    更新日期:2017-11-27 00:00:00

  • Mapping quantitative trait loci in a non-equilibrium population.

    abstract::The genetic control of a complex trait can be studied by testing and mapping the genotypes of the underlying quantitative trait loci (QTLs) through their associations with observable marker genotypes. All existing statistical methods for QTL mapping assume an equilibrium population, allowing marker-QTL associations to...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1578

    authors: Wu S,Yang J,Wu R

    更新日期:2010-01-01 00:00:00

  • BayesMendel: an R environment for Mendelian risk prediction.

    abstract::Several important syndromes are caused by deleterious germline mutations of individual genes. In both clinical and research applications it is useful to evaluate the probability that an individual carries an inherited genetic variant of these genes, and to predict the risk of disease for that individual, using informa...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1063

    authors: Chen S,Wang W,Broman KW,Katki HA,Parmigiani G

    更新日期:2004-01-01 00:00:00

  • Approximating the variance of the conditional probability of the state of a hidden Markov model.

    abstract::In a hidden Markov model, one "estimates" the state of the hidden Markov chain at t by computing via the forwards-backwards algorithm the conditional distribution of the state vector given the observed data. The covariance matrix of this conditional distribution measures the information lost by failure to observe dire...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章,评审

    doi:10.2202/1544-6115.1296

    authors: Siegmund DO,Yakir B

    更新日期:2007-01-01 00:00:00

  • Principal component discriminant analysis.

    abstract::The approach adopted involved two-stages. First the 11205 measurements in the mass spectrometry data were reduced to 14 scores by a principal component analysis of the centered but otherwise untreated and unscaled data matrix. Then a linear classifier was derived by linear discriminant analysis using these 14 scores a...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1350

    authors: Fearn T

    更新日期:2008-01-01 00:00:00

  • Detecting outlier samples in microarray data.

    abstract::In this paper, we address the problem of detecting outlier samples with highly different expression patterns in microarray data. Although outliers are not common, they appear even in widely used benchmark data sets and can negatively affect microarray data analysis. It is important to identify outliers in order to exp...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1426

    authors: Shieh AD,Hung YS

    更新日期:2009-01-01 00:00:00

  • A test for detecting differential indirect trans effects between two groups of samples.

    abstract::Integrative analysis of copy number and gene expression data can help in understanding the cis and trans effect of copy number aberrations on transcription levels of genes involved in a pathway. To analyse how these copy number mediated gene-gene interactions differ between groups of samples we propose a new method, n...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2017-0058

    authors: Chaturvedi N,Menezes RX,Goeman JJ,Wieringen WV

    更新日期:2018-07-31 00:00:00

  • Genetic linkage analysis in the presence of germline mosaicism.

    abstract::Germline mosaicism is a genetic condition in which some germ cells of an individual contain a mutation. This condition violates the assumptions underlying classic genetic analysis and may lead to failure of such analysis. In this work we extend the statistical model used for genetic linkage analysis in order to incorp...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1709

    authors: Weissbrod O,Geiger D

    更新日期:2011-10-04 00:00:00

  • Second order optimization for the inference of gene regulatory pathways.

    abstract::With the increasing availability of experimental data on gene interactions, modeling of gene regulatory pathways has gained special attention. Gradient descent algorithms have been widely used for regression and classification applications. Unfortunately, results obtained after training a model by gradient descent are...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2012-0021

    authors: Das M,Murthy CA,De RK

    更新日期:2014-02-01 00:00:00

  • On an extended interpretation of linkage disequilibrium in genetic case-control association studies.

    abstract::We are concerned with statistical inference for 2 × C × K contingency tables in the context of genetic case-control association studies. Multivariate methods based on asymptotic Gaussianity of vectors of test statistics require information about the asymptotic correlation structure among these test statistics under th...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2015-0024

    authors: Dickhaus T,Stange J,Demirhan H

    更新日期:2015-11-01 00:00:00

  • Asymptotic optimality of likelihood-based cross-validation.

    abstract::Likelihood-based cross-validation is a statistical tool for selecting a density estimate based on n i.i.d. observations from the true density among a collection of candidate density estimators. General examples are the selection of a model indexing a maximum likelihood estimator, and the selection of a bandwidth index...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1036

    authors: van der Laan MJ,Dudoit S,Keles S

    更新日期:2004-01-01 00:00:00

  • Genetic association test based on principal component analysis.

    abstract::Many gene- and pathway-based association tests have been proposed in the literature. Among them, the SKAT is widely used, especially for rare variants association studies. In this paper, we investigate the connection between SKAT and a principal component analysis. This investigation leads to a procedure that encompas...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2016-0061

    authors: Chen Z,Han S,Wang K

    更新日期:2017-07-26 00:00:00

  • Model selection based on FDR-thresholding optimizing the area under the ROC-curve.

    abstract::We evaluate variable selection by multiple tests controlling the false discovery rate (FDR) to build a linear score for prediction of clinical outcome in high-dimensional data. Quality of prediction is assessed by the receiver operating characteristic curve (ROC) for prediction in independent patients. Thus we try to ...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1462

    authors: Graf AC,Bauer P

    更新日期:2009-01-01 00:00:00

  • Polyunphased: an extension to polytomous outcomes of the Unphased package for family-based genetic association analysis.

    abstract::Polytomous phenotypes arise when a disease has multiple subtypes or when two dichotomous phenotypes are analyzed simultaneously. Few software programs offer the option to analyze such phenotypes in family studies, and none implements conditional polytomous logistic regression for within-family analysis robust to popul...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2016-0035

    authors: Bureau A,Croteau J

    更新日期:2017-03-01 00:00:00

  • Predicting protein concentrations with ELISA microarray assays, monotonic splines and Monte Carlo simulation.

    abstract::Making sound proteomic inferences using ELISA microarray assay requires both an accurate prediction of protein concentration and a credible estimate of its error. We present a method using monotonic spline statistical models (MS), penalized constrained least squares fitting (PCLS) and Monte Carlo simulation (MC) to pr...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1364

    authors: Daly DS,Anderson KK,White AM,Gonzalez RM,Varnum SM,Zangar RC

    更新日期:2008-01-01 00:00:00

  • Sparse inverse of covariance matrix of QTL effects with incomplete marker data.

    abstract::Gametic models for fitting breeding values at QTL as random effects in outbred populations have become popular because they require few assumptions about the number and distribution of QTL alleles segregating. The covariance matrix of the gametic effects has an inverse that is sparse and can be constructed rapidly by ...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1048

    authors: Thallman RM,Hanford KJ,Kachman SD,Van Vleck LD

    更新日期:2004-01-01 00:00:00

  • LCox: a tool for selecting genes related to survival outcomes using longitudinal gene expression data.

    abstract::Longitudinal genomics data and survival outcome are common in biomedical studies, where the genomics data are often of high dimension. It is of great interest to select informative longitudinal biomarkers (e.g. genes) related to the survival outcome. In this paper, we develop a computationally efficient tool, LCox, fo...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2017-0060

    authors: Sun J,Herazo-Maya JD,Wang JL,Kaminski N,Zhao H

    更新日期:2019-02-13 00:00:00