Knowledge discovery by attribute-oriented approach under directed acyclic concept graph (dacg)
Junping Sun, Wenyi Bi
Abstract
Junping Sun, Wenyi Bi
Abstract
Knowledge discovery in databases (KDD) is an active and promising research area with potentially high payoffs in business and scientific applications. The great challenge of knowledge discovery in databases is to process large quantities of raw data automatically, to identify the most significant and meaningful patterns, and to present this knowledge in an appropriate form for decision making and other purposes. In previous researches, Attribute-Oriented Induction, implemented artificial intelligence, “learning from examples” paradigm. This method integrates traditional database operations to extract rules from database systems. The key techniques in attribute-oriented induction are attribute generalization and undesirable attribute removal. Attribute generalization is implemented by replacing a low-level concept with its corresponding high level concept. The core part of this approach is a concept hierarchy, which is a linear tree schema built on each individual and independent domain (attribute), to control concept generalization. Because such linear structure of a concept hierarchy represents the concepts that are confined to each independent domain, this topology leads to a learning process without the capability of conditional concept generalization. Therefore, it is unable to extract rich knowledge implied in different directions of non-linear concept scheme. Although some recent improvements have extended to the basic attribute-oriented induction (BAOI) approach, they have some shortcomings. For example, rule-based attribute-oriented induction has to invoke a backtracking algorithm to tackle information loss problem , whereas path id generalization has to transform each data values (at a great cost) in databases into its corresponding path id in order to perform generalization on the path id relation instead.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Knowledge discovery in databases (KDD) is an active and promising research area with potentially high payoffs in business and scientific applications. The great challenge of knowledge discovery in databases is to process large quantities of raw data automatically, to identify the most significant and meaningful patterns, and to present this knowledge in an appropriate form for decision making and other purposes. In previous researches, Attribute-Oriented Induction, implemented artificial intelligence, “learning from examples” paradigm. This method integrates traditional database operations to extract rules from database systems. The key techniques in attribute-oriented induction are attribute generalization and undesirable attribute removal. Attribute generalization is implemented by replacing a low-level concept with its corresponding high level concept. The core part of this approach is a concept hierarchy, which is a linear tree schema built on each individual and independent domain (attribute), to control concept generalization. Because such linear structure of a concept hierarchy represents the concepts that are confined to each independent domain, this topology leads to a learning process without the capability of conditional concept generalization. Therefore, it is unable to extract rich knowledge implied in different directions of non-linear concept scheme. Although some recent improvements have extended to the basic attribute-oriented induction (BAOI) approach, they have some shortcomings. For example, rule-based attribute-oriented induction has to invoke a backtracking algorithm to tackle information loss problem , whereas path id generalization has to transform each data values (at a great cost) in databases into its corresponding path id in order to perform generalization on the path id relation instead.
Key concepts: Computer science, Generalization, Knowledge extraction, Data mining, Formal concept analysis, Attribute domain, Schema (genetic algorithms), Theoretical computer science