Towards Planning Scientific Experiments through Declarative Model Discovery in Provenance Data
Mateus Ferreira Silva, Fernanda Baião, Kate Cerqueira Revoredo
Abstract
Mateus Ferreira Silva, Fernanda Baião, Kate Cerqueira Revoredo
Abstract
Data provenance is the process of managing a collection of metadata that catalogs the origin and history of data. In scientific workflows, this metadata assists scientists and domain specialists in several tasks, including the reproduction of scientific experiments and planning of new scenarios to be experimented. However, the amount of provenance data generated from scientific workflow executions can grow through time, becoming infeasible for scientists to manually evaluate them. Thus, mechanisms for automatically extracting knowledge from provenance data and presenting them to the user are demanding. Due to the diversity and flexibility inherent to scientific experimentation scenarios, declarative models are potentially adequate. In this work, we propose to apply techniques for learning a declarative model from provenance data generated by scientific workflows, from which the domain specialist will be able to plan future scenarios for his/her scientific experiment. The proposed solution is illustrated in a case study of a scientific experiment on ontology matching, which is a data-intensive strategy that is required to solve the problem of information integration in several areas of knowledge.
OpenAlex reports 4 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Data provenance is the process of managing a collection of metadata that catalogs the origin and history of data. In scientific workflows, this metadata assists scientists and domain specialists in several tasks, including the reproduction of scientific experiments and planning of new scenarios to be experimented. However, the amount of provenance data generated from scientific workflow executions can grow through time, becoming infeasible for scientists to manually evaluate them. Thus, mechanisms for automatically extracting knowledge from provenance data and presenting them to the user are demanding. Due to the diversity and flexibility inherent to scientific experimentation scenarios, declarative models are potentially adequate. In this work, we propose to apply techniques for learning a declarative model from provenance data generated by scientific workflows, from which the domain specialist will be able to plan future scenarios for his/her scientific experiment. The proposed solution is illustrated in a case study of a scientific experiment on ontology matching, which is a data-intensive strategy that is required to solve the problem of information integration in several areas of knowledge.
Key concepts: Workflow, Computer science, Metadata, Ontology, Data science, Domain (mathematical analysis), Process (computing), Information retrieval