Predicting the bioconcentration factor through a conformation-independent QSPR study
José F. Aranda, Daniel E. Bacelo, M. S. Leguizamón Aparicio, Marco A. Ocsachoque, Eduardo A. Castro, Pablo R. Duchowicz
Abstract
Open-access reader
José F. Aranda, Daniel E. Bacelo, M. S. Leguizamón Aparicio, Marco A. Ocsachoque, Eduardo A. Castro, Pablo R. Duchowicz
Abstract
Open-access reader
The ANTARES dataset is a large collection of known and verified experimental bioconcentration factor data, involving 851 highly heterogeneous compounds from which 159 are pesticides. The BCF ANTARES data were used to derive a conformation-independent QSPR model. A large set of 27,017 molecular descriptors was explored, with the main intention of capturing the most relevant structural characteristics affecting the studied property. The structural descriptors were derived with different freeware tools, such as PaDEL, Epi Suite, CORAL, Mold2, RECON, and QuBiLs-MAS, and so it was interesting to find out the way that the different descriptor tools complemented each other in order to improve the statistical quality of the established QSPR. The best multivariable linear regression models were found with the Replacement Method variable sub-set selection technique. The proposed QSPR model improves previous reported models of the bioconcentration factor in the present dataset.
OpenAlex reports 24 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The ANTARES dataset is a large collection of known and verified experimental bioconcentration factor data, involving 851 highly heterogeneous compounds from which 159 are pesticides. The BCF ANTARES data were used to derive a conformation-independent QSPR model. A large set of 27,017 molecular descriptors was explored, with the main intention of capturing the most relevant structural characteristics affecting the studied property. The structural descriptors were derived with different freeware tools, such as PaDEL, Epi Suite, CORAL, Mold2, RECON, and QuBiLs-MAS, and so it was interesting to find out the way that the different descriptor tools complemented each other in order to improve the statistical quality of the established QSPR. The best multivariable linear regression models were found with the Replacement Method variable sub-set selection technique. The proposed QSPR model improves previous reported models of the bioconcentration factor in the present dataset.
Key concepts: Quantitative structure–activity relationship, Bioconcentration, Molecular descriptor, Set (abstract data type), Linear regression, Feature selection, Variable (mathematics), Chemistry