This paper presents a complete inductive learning system that aims to produce comprehensible theories for XML document classifications. The knowledge representation method is based on a higher-order logic formalism which is particularly suitable for structured-data learning systems. A systematic way of generating predicates is also given. The learning algorithm of the system is a modified standard decision-tree learning algorithm driven by predicate/recall breakeven point. Experimental results on XML version of Reuters dataset show that this system is able to produce comprehensible theories with high precision/recall breakeven point values.
XML plays a key role in service-oriented computing (SOC). We envisage an environment where knowledge workers are developing semantic correspondences from XML schemas to a target ontology. We would like to use DL reasoning services within this environment to assist the knowledge worker in the following three tasks: establishing the correctness of structural correspondences; the implications of correspondences; and finally, search and selection of schemas with semantic and structural constraints. A knowledge representation method for XML schemas using a description logic language is presented. With this representation method, it is now possible to use the underlying DL reasoning services to achieve our three tasks: correspondence correctness checking through DL model validation; correspondence implication through DL model completion; and finally, semantic and structural search and selection through DL model query.