Abstract:
:In cellular biology, node-and-edge graph or "network" data collection often uses bait-prey technologies such as co-immunoprecipitation (CoIP). Bait-prey technologies assay relationships or "interactions" between protein pairs, with CoIP specifically measuring protein complex co-membership. Analyses of CoIP data frequently focus on estimating protein complex membership. Due to budgetary and other constraints, exhaustive assay of the entire network using CoIP is not always possible. We describe a stratified sampling scheme to select baits for CoIP experiments when protein complex estimation is the main goal. Expanding upon the classic framework in which nodes represent proteins and edges represent pairwise interactions, we define generalized nodes as sets of adjacent nodes with identical adjacency outside the set and use these as strata from which to select the next set of baits. Strata are redefined at each round of sampling to incorporate accumulating data. This scheme maintains user-specified quality thresholds for protein complex estimates and, relative to simple random sampling, leads to a marked increase in the number of correctly estimated complexes at each round of sampling. The R package seqSample contains all source code and is available at http://vault.northwestern.edu/~dms877/Rpacks/.
journal_name
Stat Appl Genet Mol Biolauthors
Scholtens DM,Spencer BDdoi
10.1515/sagmb-2015-0007subject
Has Abstractpub_date
2015-08-01 00:00:00pages
391-411issue
4eissn
2194-6302issn
1544-6115pii
/j/sagmb.2015.14.issue-4/sagmb-2015-0007/sagmb-201journal_volume
14pub_type
杂志文章abstract::We address a potential shortcoming of three probabilistic models for detecting interspecific recombination in DNA sequence alignments: the multiple change-point model (MCP) of Suchard et al. (2003), the dual multiple change-point model (DMCP) of Minin et al. (2005), and the phylogenetic factorial hidden Markov model (...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1399
更新日期:2008-01-01 00:00:00
abstract::Likelihood-based cross-validation is a statistical tool for selecting a density estimate based on n i.i.d. observations from the true density among a collection of candidate density estimators. General examples are the selection of a model indexing a maximum likelihood estimator, and the selection of a bandwidth index...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1036
更新日期:2004-01-01 00:00:00
abstract::In this paper, we address the problem of detecting outlier samples with highly different expression patterns in microarray data. Although outliers are not common, they appear even in widely used benchmark data sets and can negatively affect microarray data analysis. It is important to identify outliers in order to exp...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1426
更新日期:2009-01-01 00:00:00
abstract::With the increasing availability of experimental data on gene interactions, modeling of gene regulatory pathways has gained special attention. Gradient descent algorithms have been widely used for regression and classification applications. Unfortunately, results obtained after training a model by gradient descent are...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2012-0021
更新日期:2014-02-01 00:00:00
abstract::Germline mosaicism is a genetic condition in which some germ cells of an individual contain a mutation. This condition violates the assumptions underlying classic genetic analysis and may lead to failure of such analysis. In this work we extend the statistical model used for genetic linkage analysis in order to incorp...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1709
更新日期:2011-10-04 00:00:00
abstract::We are concerned with statistical inference for 2 × C × K contingency tables in the context of genetic case-control association studies. Multivariate methods based on asymptotic Gaussianity of vectors of test statistics require information about the asymptotic correlation structure among these test statistics under th...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2015-0024
更新日期:2015-11-01 00:00:00
abstract::We present an approach to association studies involving a dozen or so ;response' variables and a few hundred ;explanatory' variables which emphasizes transparency, simplicity, and protection against spurious results. The methods proposed are largely non-parametric, and they are systematically rounded-off by the Benjam...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1420
更新日期:2009-01-01 00:00:00
abstract::Mass spectrometry is an important high-throughput technique for profiling small molecular compounds in biological samples and is widely used to identify potential diagnostic and prognostic compounds associated with disease. Commonly, this data generated by mass spectrometry has many missing values resulting when a com...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2013-0021
更新日期:2013-12-01 00:00:00
abstract::In candidate gene association studies, usually several elementary hypotheses are tested simultaneously using one particular set of data. The data normally consist of partly correlated SNP information. Every SNP can be tested for association with the disease, e.g., using the Cochran-Armitage test for trend. To account ...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1729
更新日期:2011-01-01 00:00:00
abstract::In a hidden Markov model, one "estimates" the state of the hidden Markov chain at t by computing via the forwards-backwards algorithm the conditional distribution of the state vector given the observed data. The covariance matrix of this conditional distribution measures the information lost by failure to observe dire...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章,评审
doi:10.2202/1544-6115.1296
更新日期:2007-01-01 00:00:00
abstract::Usually, a pedigree is sampled and included in the sample that is analyzed after following a predefined non-random sampling design comprising several specific procedures. To obtain a pedigree analysis result free from the bias caused by the sampling procedures, a correction is applied to the pedigree likelihood. The s...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1003
更新日期:2003-01-01 00:00:00
abstract::Combining correlated p-values from multiple hypothesis testing is a most frequently used method for integrating information in genetic and genomic data analysis. However, most existing methods for combining independent p-values from individual component problems into a single unified p-value are unsuitable for the cor...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2019-0057
更新日期:2020-11-06 00:00:00
abstract::Gene microarray technology is often used to compare the expression of thousand of genes in two different cell lines. Typically, one does not expect measurable changes in transcription amounts for a large number of genes; furthermore, the noise level of array experiments is rather high in relation to the available numb...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1132
更新日期:2005-01-01 00:00:00
abstract::Locus heterogeneity is one of the most important issues in gene mapping and can cause significant reductions in statistical power for gene mapping, yet no research to date has provided power and sample size calculations for family-based association methods in the presence of locus heterogeneity. The purpose of this re...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1501
更新日期:2009-01-01 00:00:00
abstract::Accurately measuring epigenetic marks such as 5-methylcytosine (5-mC) and 5-hydroxymethylcytosine (5-hmC) at the single-nucleotide level, requires combining data from DNA processing methods including traditional (BS), oxidative (oxBS) or Tet-Assisted (TAB) bisulfite conversion. We introduce the R package MLML2R, which...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2018-0031
更新日期:2019-01-17 00:00:00
abstract::There has been increasing interest in predicting patients' survival after therapy by investigating gene expression microarray data. In the regression and classification models with high-dimensional genomic data, boosting has been successfully applied to build accurate predictive models and conduct variable selection s...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1550
更新日期:2010-01-01 00:00:00
abstract::Recent analytic and technological breakthroughs have set the stage for genome-wide linkage disequilibrium studies to map disease-susceptibility variants. This paper discusses a probabilistic methodology for making disease-mapping inferences in large-scale case-control genetic studies. The semi-Bayesian approach promot...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1168
更新日期:2005-01-01 00:00:00
abstract::The genetic control of a complex trait can be studied by testing and mapping the genotypes of the underlying quantitative trait loci (QTLs) through their associations with observable marker genotypes. All existing statistical methods for QTL mapping assume an equilibrium population, allowing marker-QTL associations to...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1578
更新日期:2010-01-01 00:00:00
abstract::We evaluate variable selection by multiple tests controlling the false discovery rate (FDR) to build a linear score for prediction of clinical outcome in high-dimensional data. Quality of prediction is assessed by the receiver operating characteristic curve (ROC) for prediction in independent patients. Thus we try to ...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1462
更新日期:2009-01-01 00:00:00
abstract::Multi-color optical mapping is a new technique being developed to obtain detailed physical maps (indicating relative positions of various recognition sites) of DNA molecules. We consider a study design in which the data consist of noisy observations of multiple copies of a DNA molecule marked with colors at recognitio...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1266
更新日期:2007-01-01 00:00:00
abstract::The Dirichlet Process (DP) mixture model has become a popular choice for model-based clustering, largely because it allows the number of clusters to be inferred. The sequential updating and greedy search (SUGS) algorithm (Wang & Dunson, 2011) was proposed as a fast method for performing approximate Bayesian inference ...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2018-0065
更新日期:2019-12-12 00:00:00
abstract::We present a Bayesian hierarchical model for detecting differentially expressed genes using a mixture prior on the parameters representing differential effects. We formulate an easily interpretable 3-component mixture to classify genes as over-expressed, under-expressed and non-differentially expressed, and model gene...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1314
更新日期:2007-01-01 00:00:00
abstract::We present a weighted-LASSO method to infer the parameters of a first-order vector auto-regressive model that describes time course expression data generated by directed gene-to-gene regulation networks. These networks are assumed to own prior internal structures of connectivity which drive the inference method. This ...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1519
更新日期:2010-01-01 00:00:00
abstract::Multiple testing procedures are commonly used in gene expression studies for the detection of differential expression, where typically thousands of genes are measured over at least two experimental conditions. Given the need for powerful testing procedures, and the attendant danger of false positives in multiple testi...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1302
更新日期:2007-01-01 00:00:00
abstract::In recent years, alignment-free methods have been widely applied in comparing genome sequences, as these methods compute efficiently and provide desirable phylogenetic analysis results. These methods have been successfully combined with hierarchical clustering methods for finding phylogenetic trees. However, it may no...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2018-0045
更新日期:2019-02-15 00:00:00
abstract::Reproducibility of disease signatures and clinical biomarkers in multi-omics disease analysis has been a key challenge due to a multitude of factors. The heterogeneity of the limited sample, various biological factors such as environmental confounders, and the inherent experimental and technical noises, compounded wit...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2018-0039
更新日期:2019-05-11 00:00:00
abstract::Polytomous phenotypes arise when a disease has multiple subtypes or when two dichotomous phenotypes are analyzed simultaneously. Few software programs offer the option to analyze such phenotypes in family studies, and none implements conditional polytomous logistic regression for within-family analysis robust to popul...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2016-0035
更新日期:2017-03-01 00:00:00
abstract::The problem of finding periodically expressed genes from time course microarray experiments is at the center of numerous efforts to identify the molecular components of biological clocks. We present a new approach to this problem based on the cyclohedron test, which is a rank test inspired by recent advances in algebr...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1286
更新日期:2007-01-01 00:00:00
abstract::Multiple branching trees have been used to model the acquisition of HIV drug resistance mutations, and several different algorithms have been developed to construct the tree set that best describes the data. These algorithms have mainly focused on the structure of the tree set. The focal point of this paper is estimat...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.2202/1544-6115.1324
更新日期:2008-01-01 00:00:00
abstract::Integrative analysis of copy number and gene expression data can help in understanding the cis and trans effect of copy number aberrations on transcription levels of genes involved in a pathway. To analyse how these copy number mediated gene-gene interactions differ between groups of samples we propose a new method, n...
journal_title:Statistical applications in genetics and molecular biology
pub_type: 杂志文章
doi:10.1515/sagmb-2017-0058
更新日期:2018-07-31 00:00:00