Consensus QSAR models: do the benefits outweigh the complexity?

Abstract:

:This study has assessed the use of consensus regression, as compared to single multiple linear regression, models for the development of quantitative structure-activity relationships (QSARs). To provide a comparison, four data sets of varying size and complexity were analyzed: silastic membrane flux, toxicity of phenols to Tetrahymena pyriformis, acute toxicity to the fathead minnow and flash point. For each data set, a genetic algorithm was used to develop a model population and the performance of consensus models was compared to that of the best single model. Two consensus models were developed, one using the top 10 models, and the other using a subset of models chosen to provide maximal coverage of model space. The results highlight the ability of the genetic algorithm to develop predictive models from a large descriptor pool. However, the consensus models were shown to offer no significant improvements over single regression models, which are as statistically robust as the equivalent consensus models. Consensus models developed from a selection of the best QSARs were shown not to be superior to a selection of diverse in "model space" QSARs. For the data sets analyzed in this study, and in light of the Organization for Economic Cooperation and Development principles for the validation of QSARs, the increase in model complexity when using consensus models does not seem warranted given the minimal improvement in model statistics.

journal_name

J Chem Inf Model

authors

Hewitt M,Cronin MT,Madden JC,Rowe PH,Johnson C,Obi A,Enoch SJ

doi

10.1021/ci700016d

subject

Has Abstract

pub_date

2007-07-01 00:00:00

pages

1460-8

issue

4

eissn

1549-9596

issn

1549-960X

journal_volume

47

pub_type

杂志文章
  • Flux (1): a virtual synthesis scheme for fragment-based de novo design.

    abstract::It is demonstrated that the fragmentation of druglike molecules by applying simplistic pseudo-retrosynthesis results in a stock of chemically meaningful building blocks for de novo molecule generation. A stochastic search algorithm in conjunction with ligand-based similarity scoring (Flux: fragment-based ligand builde...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci0503560

    authors: Fechner U,Schneider G

    更新日期:2006-03-01 00:00:00

  • Accurate Estimation of pKb Values for Amino Groups from Surface Electrostatic Potential (VS,min) Calculations: The Isoelectric Points of Amino Acids as a Case Study.

    abstract::Theoretical calculation of equilibrium dissociation constants is a very computationally demanding and time-consuming process since it requires an extremely accurate computation of the solvation free energy changes for each of the species involved. By correlating the minimum surface electrostatic potential (VS,min) on ...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.9b01173

    authors: Sandoval-Lira J,Mondragón-Solórzano G,Lugo-Fuentes LI,Barroso-Flores J

    更新日期:2020-03-23 00:00:00

  • Universal Activation Index for Class A GPCRs.

    abstract::An index of the activation of Class A G-protein-coupled receptors (GPCRs) has been trained using interhelix distances from a series of microsecond molecular-dynamics simulations and tested for 268 published X-ray structures. In a three-class model that includes intermediate structures, 63% of the active structures are...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.9b00604

    authors: Ibrahim P,Wifling D,Clark T

    更新日期:2019-09-23 00:00:00

  • Prediction of the Favorable Hydration Sites in a Protein Binding Pocket and Its Application to Scoring Function Formulation.

    abstract::The important role of water molecules in protein-ligand binding energetics has attracted wide attention in recent years. A range of computational methods has been developed to predict the favorable locations of water molecules in a protein binding pocket. Most of the current methods are based on extensive molecular dy...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.9b00619

    authors: Li Y,Gao Y,Holloway MK,Wang R

    更新日期:2020-09-28 00:00:00

  • Spatial sign preprocessing: a simple way to impart moderate robustness to multivariate estimators.

    abstract::The spatial sign is a multivariate extension of the concept of sign. Recently multivariate estimators of covariance structures based on spatial signs have been examined by various authors. These new estimators are found to be robust to outlying observations. From a computational point of view, estimators based on spat...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci050498u

    authors: Serneels S,De Nolf E,Van Espen PJ

    更新日期:2006-05-01 00:00:00

  • Discovery of New SIRT2 Inhibitors by Utilizing a Consensus Docking/Scoring Strategy and Structure-Activity Relationship Analysis.

    abstract::SIRT2, which is a NAD+ (nicotinamide adenine dinucleotide) dependent deacetylase, has been demonstrated to play an important role in the occurrence and development of a variety of diseases such as cancer, ischemia-reperfusion, and neurodegenerative diseases. Small molecule inhibitors of SIRT2 are thought to be potenti...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.6b00714

    authors: Huang S,Song C,Wang X,Zhang G,Wang Y,Jiang X,Sun Q,Huang L,Xiang R,Hu Y,Li L,Yang S

    更新日期:2017-04-24 00:00:00

  • OPUS-Rota3: Improving Protein Side-Chain Modeling by Deep Neural Networks and Ensemble Methods.

    abstract::Side-chain modeling is critical for protein structure prediction since the uniqueness of the protein structure is largely determined by its side-chain packing conformation. In this paper, differing from most approaches that rely on rotamer library sampling, we first propose a novel side-chain rotamer prediction method...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.0c00951

    authors: Xu G,Wang Q,Ma J

    更新日期:2020-12-28 00:00:00

  • Reaction site mapping of xenobiotic biotransformations.

    abstract::Predictive metabolism methods can be used in drug discovery projects to enhance the understanding of structure-metabolism relationships. The present study uses data mining methods to exploit biotransformation data that have been recorded in the MDL Metabolite database. Reacting center fingerprints were derived from a ...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci600376q

    authors: Boyer S,Arnby CH,Carlsson L,Smith J,Stein V,Glen RC

    更新日期:2007-03-01 00:00:00

  • Target-independent prediction of drug synergies using only drug lipophilicity.

    abstract::Physicochemical properties of compounds have been instrumental in selecting lead compounds with increased drug-likeness. However, the relationship between physicochemical properties of constituent drugs and the tendency to exhibit drug interaction has not been systematically studied. We assembled physicochemical descr...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci500276x

    authors: Yilancioglu K,Weinstein ZB,Meydan C,Akhmetov A,Toprak I,Durmaz A,Iossifov I,Kazan H,Roth FP,Cokol M

    更新日期:2014-08-25 00:00:00

  • Modeling oral rat chronic toxicity.

    abstract::The chronic toxicity is fundamental for toxicological risk assessment, but its correlation with the chemical structures has been studied only little. This is partly due to the complexity of such an experimental test that embraces a plethora of different biological effects and mechanisms of action, making (Q)SAR studie...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci8001974

    authors: Mazzatorta P,Estevez MD,Coulet M,Schilter B

    更新日期:2008-10-01 00:00:00

  • Rigorous Computational Study Reveals What Docking Overlooks: Double Trouble from Membrane Association in Protein Kinase C Modulators.

    abstract::Increasing protein kinase C (PKC) activity is of potential therapeutic value. Its activation involves an interaction between the C1 domain and diacylglycerol (DAG) at intracellular membrane surfaces; DAG mimetics hold promise as new drugs. We previously developed the isophthalate derivative HMI-1a3, an effective but h...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.0c00624

    authors: Lautala S,Provenzani R,Koivuniemi A,Kulig W,Talman V,Róg T,Tuominen RK,Yli-Kauhaluoma J,Bunker A

    更新日期:2020-11-23 00:00:00

  • Development of an informatics platform for therapeutic protein and peptide analytics.

    abstract::The momentum gained by research on biologics has not been met yet with equal thrust on the informatics side. There is a noticeable lack of software for data management that empowers the bench scientists working on the development of biologic therapeutics. SARvision|Biologics is a tool to analyze data associated with b...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci400333x

    authors: Hansen MR,Villar HO,Feyfant E

    更新日期:2013-10-28 00:00:00

  • Efficient Corrections for DFT Noncovalent Interactions Based on Ensemble Learning Models.

    abstract::Machine learning has exhibited powerful capabilities in many areas. However, machine learning models are mostly database dependent, requiring a new model if the database changes. Therefore, a universal model is highly desired to accommodate the widest variety of databases. Fortunately, this universality may be achieve...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.8b00878

    authors: Li W,Miao W,Cui J,Fang C,Su S,Li H,Hu L,Lu Y,Chen G

    更新日期:2019-05-28 00:00:00

  • Exploring Tunable Hyperparameters for Deep Neural Networks with Industrial ADME Data Sets.

    abstract::Deep learning has drawn significant attention in different areas including drug discovery. It has been proposed that it could outperform other machine learning algorithms, especially with big data sets. In the field of pharmaceutical industry, machine learning models are built to understand quantitative structure-acti...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.8b00671

    authors: Zhou Y,Cahya S,Combs SA,Nicolaou CA,Wang J,Desai PV,Shen J

    更新日期:2019-03-25 00:00:00

  • The valence state combination model: a generic framework for handling tautomers and protonation states.

    abstract::The consistent handling of molecules is probably the most basic and important requirement in the field of cheminformatics. Reliable results can only be obtained if the underlying calculations are independent of the specific way molecules are represented in the input data. However, ensuring consistency is a complex tas...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci400724v

    authors: Urbaczek S,Kolodzik A,Rarey M

    更新日期:2014-03-24 00:00:00

  • Adaptive configuring of radial basis function network by hybrid particle swarm algorithm for QSAR studies of organic compounds.

    abstract::The configuring of a radial basis function network (RBFN) consists of selecting the network parameters (centers and widths in RBF units and weights between the hidden and output layers) and network architecture. The issues of suboptimum and overfitting, however, often occur in RBFN configuring. This paper presented a ...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci600218d

    authors: Zhou YP,Jiang JH,Lin WQ,Zou HY,Wu HL,Shen GL,Yu RQ

    更新日期:2006-11-01 00:00:00

  • Determination of partition coefficient of spin probe between different lipid membrane phases.

    abstract::Model lipid membranes made from binary mixtures of dimyristoylphosphatidylcholine/dipalmitoylphosphatidylcholine (DMPC/DPPC) and dimyristoylphosphatidylcholine/cholesterol (DMPC/Chol) exhibit coexistence of diverse lipid phases at appropriate temperature and composition. Since lipids in different phases show different...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci0501793

    authors: Arsov Z,Strancar J

    更新日期:2005-11-01 00:00:00

  • Effects of Ligand Environment in Zr(IV) Assisted Peptide Hydrolysis.

    abstract::In this DFT study, activities of 11 different N2O4, N2O3, and NO2 core containing Zr(IV) complexes, 4,13-diaza-18-crown-6 (I'N2O4), 1,4,10-trioxa-7,13-diazacyclopentadecane (I'N2O3), and 2-(2-methoxy)ethanol (I'NO2), respectively, and their analogues in peptide hydrolysis have been investigated. Based on the experimen...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.6b00781

    authors: Zhang T,Sharma G,Paul TJ,Hoffmann Z,Prabhakar R

    更新日期:2017-05-22 00:00:00

  • Symplectic molecular dynamics simulations on specially designed parallel computers.

    abstract::We have developed a computer program for molecular dynamics (MD) simulation that implements the Split Integration Symplectic Method (SISM) and is designed to run on specialized parallel computers. The MD integration is performed by the SISM, which analytically treats high-frequency vibrational motion and thus enables ...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci050216q

    authors: Borstnik U,Janezic D

    更新日期:2005-11-01 00:00:00

  • The normal-mode entropy in the MM/GBSA method: effect of system truncation, buffer region, and dielectric constant.

    abstract::We have performed a systematic study of the entropy term in the MM/GBSA (molecular mechanics combined with generalized Born and surface-area solvation) approach to calculate ligand-binding affinities. The entropies are calculated by a normal-mode analysis of harmonic frequencies from minimized snapshots of molecular d...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci3001919

    authors: Genheden S,Kuhn O,Mikulskis P,Hoffmann D,Ryde U

    更新日期:2012-08-27 00:00:00

  • Factors affecting d-block metal-ligand bond lengths: toward an automated library of molecular geometry for metal complexes.

    abstract::Metal-ligand (M-L) bond lengths for a range of ligands (carboxylates, chlorides, pyridines, water, tertiary phosphines, and alkenes) and a variety of metals have been retrieved from the Cambridge Structural Database, CSD. Analysis of the factors which affect M-L bond lengths (for example, ligand coordination mode, oxi...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci0500785

    authors: Harris SE,Orpen AG,Bruno IJ,Taylor R

    更新日期:2005-11-01 00:00:00

  • What do we know about C28H14 and C30H14 benzenoid hydrocarbons and their evolution to related polymer strips?

    abstract::While critically reviewing the current status of what is known about C28H14 and C30H14 benzenoid isomers, which are ubiquitous pyrolytic constituents, some new insights will be presented. Representative isomers belonging to these benzenoid hydrocarbons are at the crossroads to homologous series that extend to infinite...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci050298i

    authors: Dias JR

    更新日期:2006-03-01 00:00:00

  • Assessing the Protective Activity of a Recently Discovered Phenolic Compound against Oxidative Stress Using Computational Chemistry.

    abstract::The protection exerted by 3,5-dihydroxy-4-methoxybenzyl alcohol (DHMBA), a phenolic compound recently isolated from the Pacific oyster, against oxidative stress (OS) is investigated using the density functional theory. Our results indicate that DHMBA is an outstanding peroxyl radical scavenger, being about 15 times an...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.5b00513

    authors: Villuendas-Rey Y,Alvarez-Idaboy JR,Galano A

    更新日期:2015-12-28 00:00:00

  • Searching for New Leads To Treat Epilepsy: Target-Based Virtual Screening for the Discovery of Anticonvulsant Agents.

    abstract::The purpose of this investigation is to contribute to the development of new anticonvulsant drugs to treat patients with refractory epilepsy. We applied a virtual screening protocol that involved the search into molecular databases of new compounds and known drugs to find small molecules that interact with the open co...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.7b00721

    authors: Palestro PH,Enrique N,Goicoechea S,Villalba ML,Sabatier LL,Martin P,Milesi V,Bruno Blanch LE,Gavernet L

    更新日期:2018-07-23 00:00:00

  • Coupling of Zinc-Binding and Secondary Structure in Nonfibrillar Aβ40 Peptide Oligomerization.

    abstract::Nonfibrillar neurotoxic amyloid β (Aβ) oligomer structures are typically rich in β-sheets, which could be promoted by metal ions like Zn(2+). Here, using molecular dynamics (MD) simulations, we systematically examined combinations of Aβ40 peptide conformations and Zn(2+) binding modes to probe the effects of secondary...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.5b00063

    authors: Xu L,Shan S,Chen Y,Wang X,Nussinov R,Ma B

    更新日期:2015-06-22 00:00:00

  • Support vector regression scoring of receptor-ligand complexes for rank-ordering and virtual screening of chemical libraries.

    abstract::The community structure-activity resource (CSAR) data sets are used to develop and test a support vector machine-based scoring function in regression mode (SVR). Two scoring functions (SVR-KB and SVR-EP) are derived with the objective of reproducing the trend of the experimental binding affinities provided within the ...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci200078f

    authors: Li L,Wang B,Meroueh SO

    更新日期:2011-09-26 00:00:00

  • Determination of Structural Ensembles of Flexible Molecules in Solution from NMR Data Undergoing Spin Diffusion.

    abstract::Spin diffusion is a formidable problem when interpreting NMR data of chemical compounds. We developed a method to reconstruct the conformational ensemble of flexible molecules displaying spin diffusion, which minimizes the subjective bias in the interpretation of experimental data and which can be used routinely to ob...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.9b00259

    authors: Vasile F,Tiana G

    更新日期:2019-06-24 00:00:00

  • Direct Observation of β-Barrel Intermediates in the Self-Assembly of Toxic SOD128-38 and Absence in Nontoxic Glycine Mutants.

    abstract::Soluble low-molecular-weight oligomers formed during the early stage of amyloid aggregation are considered the major toxic species in amyloidosis. The structure-function relationship between oligomeric assemblies and the cytotoxicity in amyloid diseases are still elusive due to the heterogeneous and transient nature o...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.0c01319

    authors: Sun Y,Huang J,Duan X,Ding F

    更新日期:2021-01-14 00:00:00

  • Modeling Boronic Acid Based Fluorescent Saccharide Sensors: Computational Investigation of d-Fructose Binding to Dimethylaminomethylphenylboronic Acid.

    abstract::Designing organic saccharide sensors for use in aqueous solution is a nontrivial endeavor. Incorporation of hydrogen bonding groups on a sensor's receptor unit to target saccharides is an obvious strategy but not one that is likely to ensure analyte-receptor interactions over analyte-solvent or receptor-solvent intera...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/acs.jcim.8b00987

    authors: Kearns FL,Robart C,Kemp MT,Vankayala SL,Chapin BM,Anslyn EV,Woodcock HL,Larkin JD

    更新日期:2019-05-28 00:00:00

  • New fragment weighting scheme for the Bayesian inference network in ligand-based virtual screening.

    abstract::Many of the conventional similarity methods assume that molecular fragments that do not relate to biological activity carry the same weight as the important ones. One possible approach to this problem is to use the Bayesian inference network (BIN), which models molecules and reference structures as probabilistic infer...

    journal_title:Journal of chemical information and modeling

    pub_type: 杂志文章

    doi:10.1021/ci100232h

    authors: Abdo A,Salim N

    更新日期:2011-01-24 00:00:00