Map Metadata - Going Beyond the Obvious/Connecting the Dots
Gregory Steffens
Abstract
Gregory Steffens
Abstract
Much attention has been paid to the design and use of metadata to store data standards, study data specifications and creating the industry-standard metadata, i.e. the define.xml file. But not as much attention has been focused on the design and use of metadata to describe the transformations and derivations of data. We must expand beyond the focus on the stopping points of data flow and concentrate on the data movement from one stopping point to another. This paper describes the need for an industry standard for map metadata and presents a design that has been implemented with great success. The metadata design principles (described in earlier papers by Gregory Steffens and others) apply to map metadata, such as - 1.) Map metadata should not assume any one data standard, 2.) Welldesigned map metadata supports meta-programming. Mapping information is often just put into free-form text fields, but putting this information into structured metadata enables meta-programming of data flows. Map metadata has many advantages, besides automating data flow; such as data flow transparency; standardized ways to exchange specifications about how data flows from one structure to another, as in SDTM to ADaM to IDB; and creating a target metadatabase from a source metadatabase and map metadata, such as creating a study specification from a data standard. Map metadata defines the relationship between the source and target database at dataset, variable, row and value level. Map metadata, along with source and target metadata, can enable automation of the dataflow and create transparent define files, with transformation logic as well as database attribute descriptions. This is the next evolution of metadata and of meta-programming, leading to a true Data Transformation Engine (DTE). Automation at this level is essential to meet the efficiency and quality goals in today’s environment, which requires us to do more work with less staff and to improve quality at the same time.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Much attention has been paid to the design and use of metadata to store data standards, study data specifications and creating the industry-standard metadata, i.e. the define.xml file. But not as much attention has been focused on the design and use of metadata to describe the transformations and derivations of data. We must expand beyond the focus on the stopping points of data flow and concentrate on the data movement from one stopping point to another. This paper describes the need for an industry standard for map metadata and presents a design that has been implemented with great success. The metadata design principles (described in earlier papers by Gregory Steffens and others) apply to map metadata, such as - 1.) Map metadata should not assume any one data standard, 2.) Welldesigned map metadata supports meta-programming. Mapping information is often just put into free-form text fields, but putting this information into structured metadata enables meta-programming of data flows. Map metadata has many advantages, besides automating data flow; such as data flow transparency; standardized ways to exchange specifications about how data flows from one structure to another, as in SDTM to ADaM to IDB; and creating a target metadatabase from a source metadatabase and map metadata, such as creating a study specification from a data standard. Map metadata defines the relationship between the source and target database at dataset, variable, row and value level. Map metadata, along with source and target metadata, can enable automation of the dataflow and create transparent define files, with transformation logic as well as database attribute descriptions. This is the next evolution of metadata and of meta-programming, leading to a true Data Transformation Engine (DTE). Automation at this level is essential to meet the efficiency and quality goals in today’s environment, which requires us to do more work with less staff and to improve quality at the same time.
Key concepts: Metadata, Data element, Metadata repository, Meta Data Services, Computer science, Data mapping, Geospatial metadata, Database catalog