Transmission disequilibrium test power and sample size in the presence of locus heterogeneity.

Abstract:

:Locus heterogeneity is one of the most important issues in gene mapping and can cause significant reductions in statistical power for gene mapping, yet no research to date has provided power and sample size calculations for family-based association methods in the presence of locus heterogeneity. The purpose of this research is three-fold: (i) to provide an analytic solution to the incorporation of locus heterogeneity into power and sample size calculations for the TDT statistic; (ii) to verify our analytic solution with simulations; and (iii) to study how different factors affect sample size requirement for the TDT in the presence of locus heterogeneity. The detection of association in the presence of locus heterogeneity requires a greater sample size than in its absence. This increase is independent of the prevalence of the disease. In addition, as the proportion of families unlinked to the disease locus increases, the sample size necessary to maintain constant power increases. Finally, as the effect size of the disease locus increases, the sample size necessary to detect association decreases in the presence of locus heterogeneity. We provide freely available software that can perform these calculations.

authors

Chen C,Yang G,Buyske S,Matise T,Finch SJ,Gordon D

doi

10.2202/1544-6115.1501

subject

Has Abstract

pub_date

2009-01-01 00:00:00

pages

Article 44

eissn

2194-6302

issn

1544-6115

journal_volume

8

pub_type

杂志文章
  • Weighted-LASSO for structured network inference from time course data.

    abstract::We present a weighted-LASSO method to infer the parameters of a first-order vector auto-regressive model that describes time course expression data generated by directed gene-to-gene regulation networks. These networks are assumed to own prior internal structures of connectivity which drive the inference method. This ...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1519

    authors: Charbonnier C,Chiquet J,Ambroise C

    更新日期:2010-01-01 00:00:00

  • Genetic association test based on principal component analysis.

    abstract::Many gene- and pathway-based association tests have been proposed in the literature. Among them, the SKAT is widely used, especially for rare variants association studies. In this paper, we investigate the connection between SKAT and a principal component analysis. This investigation leads to a procedure that encompas...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2016-0061

    authors: Chen Z,Han S,Wang K

    更新日期:2017-07-26 00:00:00

  • Detecting outlier samples in microarray data.

    abstract::In this paper, we address the problem of detecting outlier samples with highly different expression patterns in microarray data. Although outliers are not common, they appear even in widely used benchmark data sets and can negatively affect microarray data analysis. It is important to identify outliers in order to exp...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1426

    authors: Shieh AD,Hung YS

    更新日期:2009-01-01 00:00:00

  • A Bayesian approach to estimation and testing in time-course microarray experiments.

    abstract::The objective of the present paper is to develop a truly functional Bayesian method specifically designed for time series microarray data. The method allows one to identify differentially expressed genes in a time-course microarray experiment, to rank them and to estimate their expression profiles. Each gene expressio...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1299

    authors: Angelini C,De Canditiis D,Mutarelli M,Pensky M

    更新日期:2007-01-01 00:00:00

  • Dimension reduction for classification with gene expression microarray data.

    abstract::An important application of gene expression microarray data is classification of biological samples or prediction of clinical and other outcomes. One necessary part of multivariate statistical analysis in such applications is dimension reduction. This paper provides a comparison study of three dimension reduction tech...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1147

    authors: Dai JJ,Lieu L,Rocke D

    更新日期:2006-01-01 00:00:00

  • Accommodating uncertainty in a tree set for function estimation.

    abstract::Multiple branching trees have been used to model the acquisition of HIV drug resistance mutations, and several different algorithms have been developed to construct the tree set that best describes the data. These algorithms have mainly focused on the structure of the tree set. The focal point of this paper is estimat...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1324

    authors: Healy BC,DeGruttola VG,Hu C

    更新日期:2008-01-01 00:00:00

  • On the operational characteristics of the Benjamini and Hochberg False Discovery Rate procedure.

    abstract::Multiple testing procedures are commonly used in gene expression studies for the detection of differential expression, where typically thousands of genes are measured over at least two experimental conditions. Given the need for powerful testing procedures, and the attendant danger of false positives in multiple testi...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1302

    authors: Green GH,Diggle PJ

    更新日期:2007-01-01 00:00:00

  • Sampling correction in pedigree analysis.

    abstract::Usually, a pedigree is sampled and included in the sample that is analyzed after following a predefined non-random sampling design comprising several specific procedures. To obtain a pedigree analysis result free from the bias caused by the sampling procedures, a correction is applied to the pedigree likelihood. The s...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1003

    authors: Ginsburg E,Malkin I,Elston RC

    更新日期:2003-01-01 00:00:00

  • Empirical bayes microarray ANOVA and grouping cell lines by equal expression levels.

    abstract::In the exploding field of gene expression techniques such as DNA microarrays, there are still few general probabilistic methods for analysis of variance. Linear models and ANOVA are heavily used tools in many other disciplines of scientific research. The usual F-statistic is unsatisfactory for microarray data, which e...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1125

    authors: Lönnstedt I,Rimini R,Nilsson P

    更新日期:2005-01-01 00:00:00

  • Combining dependent p-values by gamma distributions.

    abstract::Combining correlated p-values from multiple hypothesis testing is a most frequently used method for integrating information in genetic and genomic data analysis. However, most existing methods for combining independent p-values from individual component problems into a single unified p-value are unsuitable for the cor...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2019-0057

    authors: Chien LC

    更新日期:2020-11-06 00:00:00

  • Empirical bayes estimation of a sparse vector of gene expression changes.

    abstract::Gene microarray technology is often used to compare the expression of thousand of genes in two different cell lines. Typically, one does not expect measurable changes in transcription amounts for a large number of genes; furthermore, the noise level of array experiments is rather high in relation to the available numb...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1132

    authors: Erickson S,Sabatti C

    更新日期:2005-01-01 00:00:00

  • Semi-parametric differential expression analysis via partial mixture estimation.

    abstract::We develop an approach for microarray differential expression analysis, i.e. identifying genes whose expression levels differ between two or more groups. Current approaches to inference rely either on full parametric assumptions or on permutation-based techniques for sampling under the null distribution. In some situa...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1333

    authors: Rossell D,Guerra R,Scott C

    更新日期:2008-01-01 00:00:00

  • Combining nearest neighbor classifiers versus cross-validation selection.

    abstract::Various discriminant methods have been applied for classification of tumors based on gene expression profiles, among which the nearest neighbor (NN) method has been reported to perform relatively well. Usually cross-validation (CV) is used to select the neighbor size as well as the number of variables for the NN metho...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1054

    authors: Paik M,Yang Y

    更新日期:2004-01-01 00:00:00

  • Surveying the manifold divergence of an entire protein class for statistical clues to underlying biochemical mechanisms.

    abstract::Certain residues have no known function yet are co-conserved across distantly related protein families and diverse organisms, suggesting that they perform critical roles associated with as-yet-unidentified molecular properties and mechanisms. This raises the question of how to obtain additional clues regarding these m...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1666

    authors: Neuwald AF

    更新日期:2011-01-01 00:00:00

  • Second order optimization for the inference of gene regulatory pathways.

    abstract::With the increasing availability of experimental data on gene interactions, modeling of gene regulatory pathways has gained special attention. Gradient descent algorithms have been widely used for regression and classification applications. Unfortunately, results obtained after training a model by gradient descent are...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2012-0021

    authors: Das M,Murthy CA,De RK

    更新日期:2014-02-01 00:00:00

  • Variance and covariance heterogeneity analysis for detection of metabolites associated with cadmium exposure.

    abstract::In this study, we propose a novel statistical framework for detecting progressive changes in molecular traits as response to a pathogenic stimulus. In particular, we propose to employ Bayesian hierarchical models to analyse changes in mean level, variance and correlation of metabolic traits in relation to covariates. ...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2013-0041

    authors: Salamanca BV,Ebbels TM,Iorio MD

    更新日期:2014-04-01 00:00:00

  • Genetic linkage analysis in the presence of germline mosaicism.

    abstract::Germline mosaicism is a genetic condition in which some germ cells of an individual contain a mutation. This condition violates the assumptions underlying classic genetic analysis and may lead to failure of such analysis. In this work we extend the statistical model used for genetic linkage analysis in order to incorp...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1709

    authors: Weissbrod O,Geiger D

    更新日期:2011-10-04 00:00:00

  • The cyclohedron test for finding periodic genes in time course expression studies.

    abstract::The problem of finding periodically expressed genes from time course microarray experiments is at the center of numerous efforts to identify the molecular components of biological clocks. We present a new approach to this problem based on the cyclohedron test, which is a rank test inspired by recent advances in algebr...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1286

    authors: Morton J,Pachter L,Shiu A,Sturmfels B

    更新日期:2007-01-01 00:00:00

  • Accounting for undetected compounds in statistical analyses of mass spectrometry 'omic studies.

    abstract::Mass spectrometry is an important high-throughput technique for profiling small molecular compounds in biological samples and is widely used to identify potential diagnostic and prognostic compounds associated with disease. Commonly, this data generated by mass spectrometry has many missing values resulting when a com...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2013-0021

    authors: Taylor SL,Leiserowitz GS,Kim K

    更新日期:2013-12-01 00:00:00

  • Buckley-James boosting for survival analysis with high-dimensional biomarker data.

    abstract::There has been increasing interest in predicting patients' survival after therapy by investigating gene expression microarray data. In the regression and classification models with high-dimensional genomic data, boosting has been successfully applied to build accurate predictive models and conduct variable selection s...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1550

    authors: Wang Z,Wang CY

    更新日期:2010-01-01 00:00:00

  • Fast approximate inference for variable selection in Dirichlet process mixtures, with an application to pan-cancer proteomics.

    abstract::The Dirichlet Process (DP) mixture model has become a popular choice for model-based clustering, largely because it allows the number of clusters to be inferred. The sequential updating and greedy search (SUGS) algorithm (Wang & Dunson, 2011) was proposed as a fast method for performing approximate Bayesian inference ...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2018-0065

    authors: Crook OM,Gatto L,Kirk PDW

    更新日期:2019-12-12 00:00:00

  • Likelihood-based inference for multi-color optical mapping.

    abstract::Multi-color optical mapping is a new technique being developed to obtain detailed physical maps (indicating relative positions of various recognition sites) of DNA molecules. We consider a study design in which the data consist of noisy observations of multiple copies of a DNA molecule marked with colors at recognitio...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1266

    authors: Tong L,Mets L,McPeek MS

    更新日期:2007-01-01 00:00:00

  • Approximate maximum likelihood estimation for population genetic inference.

    abstract::In many population genetic problems, parameter estimation is obstructed by an intractable likelihood function. Therefore, approximate estimation methods have been developed, and with growing computational power, sampling-based methods became popular. However, these methods such as Approximate Bayesian Computation (ABC...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2017-0016

    authors: Bertl J,Ewing G,Kosiol C,Futschik A

    更新日期:2017-11-27 00:00:00

  • Predicting protein concentrations with ELISA microarray assays, monotonic splines and Monte Carlo simulation.

    abstract::Making sound proteomic inferences using ELISA microarray assay requires both an accurate prediction of protein concentration and a credible estimate of its error. We present a method using monotonic spline statistical models (MS), penalized constrained least squares fitting (PCLS) and Monte Carlo simulation (MC) to pr...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1364

    authors: Daly DS,Anderson KK,White AM,Gonzalez RM,Varnum SM,Zangar RC

    更新日期:2008-01-01 00:00:00

  • Addressing the shortcomings of three recent Bayesian methods for detecting interspecific recombination in DNA sequence alignments.

    abstract::We address a potential shortcoming of three probabilistic models for detecting interspecific recombination in DNA sequence alignments: the multiple change-point model (MCP) of Suchard et al. (2003), the dual multiple change-point model (DMCP) of Minin et al. (2005), and the phylogenetic factorial hidden Markov model (...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1399

    authors: Husmeier D,Mantzaris AV

    更新日期:2008-01-01 00:00:00

  • Fully Bayesian mixture model for differential gene expression: simulations and model checks.

    abstract::We present a Bayesian hierarchical model for detecting differentially expressed genes using a mixture prior on the parameters representing differential effects. We formulate an easily interpretable 3-component mixture to classify genes as over-expressed, under-expressed and non-differentially expressed, and model gene...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.2202/1544-6115.1314

    authors: Lewin A,Bochkina N,Richardson S

    更新日期:2007-01-01 00:00:00

  • On an extended interpretation of linkage disequilibrium in genetic case-control association studies.

    abstract::We are concerned with statistical inference for 2 × C × K contingency tables in the context of genetic case-control association studies. Multivariate methods based on asymptotic Gaussianity of vectors of test statistics require information about the asymptotic correlation structure among these test statistics under th...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2015-0024

    authors: Dickhaus T,Stange J,Demirhan H

    更新日期:2015-11-01 00:00:00

  • Comparison and visualisation of agreement for paired lists of rankings.

    abstract::Output from analysis of a high-throughput 'omics' experiment very often is a ranked list. One commonly encountered example is a ranked list of differentially expressed genes from a gene expression experiment, with a length of many hundreds of genes. There are numerous situations where interest is in the comparison of ...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2016-0036

    authors: Donald MR,Wilson SR

    更新日期:2017-03-01 00:00:00

  • Reproducibility of biomarker identifications from mass spectrometry proteomic data in cancer studies.

    abstract::Reproducibility of disease signatures and clinical biomarkers in multi-omics disease analysis has been a key challenge due to a multitude of factors. The heterogeneity of the limited sample, various biological factors such as environmental confounders, and the inherent experimental and technical noises, compounded wit...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2018-0039

    authors: Liang Y,Kelemen A,Kelemen A

    更新日期:2019-05-11 00:00:00

  • MLML2R: an R package for maximum likelihood estimation of DNA methylation and hydroxymethylation proportions.

    abstract::Accurately measuring epigenetic marks such as 5-methylcytosine (5-mC) and 5-hydroxymethylcytosine (5-hmC) at the single-nucleotide level, requires combining data from DNA processing methods including traditional (BS), oxidative (oxBS) or Tet-Assisted (TAB) bisulfite conversion. We introduce the R package MLML2R, which...

    journal_title:Statistical applications in genetics and molecular biology

    pub_type: 杂志文章

    doi:10.1515/sagmb-2018-0031

    authors: Kiihl SF,Martinez-Garrido MJ,Domingo-Relloso A,Bermudez J,Tellez-Plaza M

    更新日期:2019-01-17 00:00:00