NLP is mainly focused on words and sentences even if most people agree that a better understanding of text structure could help extracting knowledge. In this article we show that, in certain domains, textual approaches are relevant for NLP, taking as an example a specific task related to the medical domain : the automatic modelling of health practice guidelines. Our system, GemFrame, is capable of automatically structuring Health Practice Guidelines. We propose a strategy based on the recognition of linguistic features. The system has been validated on three complementary aspects : usefulness, performances and relevance of the method.
This paper focuses on the task that consists in automatically structuring free texts according to semantic principles, hence requiring a discourse analysis. We show that the task can be rephrased as a machine learning task in which the algorithm is supposed to take an optimal decision from the range of complex interacting constraints. The approach is implemented and evaluated taking the example of Health Practices Guidelines, medium-size documents intended to describe common practices that should be followed by physicians. Our approach outperforms previous approaches limited to sentence boundaries or requiring a lot of manual work.
Health Practice Guideliens are supposed to unify practices and propose recommendations to physicians. This paper describes GemFrame, a system capable of semi-automatically filling an XML template from free texts in the clinical domain. The XML template includes semantic information not explicitly encoded in the text (pairs of conditions and ac-tions/recommendations). Therefore, there is a need to compute the exact scope of condi-tions over text sequences expressing the re-quired actions. We present a system developped for this task. We show that it yields good performance when applied to the analysis of French practice guidelines. We conclude with a precise evaluation of the tool.
This paper describes a system capable of semi-automatically filling an XML template from free texts in the clinical domain (practice guidelines). The XML template includes semantic information not explicitly encoded in the text (pairs of conditions and actions/recommendations). Therefore, there is a need to compute the exact scope of conditions over text sequences expressing the required actions. We present in this paper the rules developed for this task. We show that the system yields good performance when applied to the analysis of French practice guidelines.
This paper describes a system capable of semi-automatically filling an XML template from free texts in the clinical domain (practice guidelines). The XML template includes semantic information not explicitly encoded in the text (pairs of conditions and actions/recommendations). Therefore, there is a need to compute the exact scope of conditions over text sequences expressing the required actions. We present a system developed for this task. We show that it yields good performance when applied to the analysis of French practice guidelines.
In this paper, we study the role of the visual organization (paragraphs, headings, lists…) for a seg- mentation task of procedural texts. We focus on a particular type of procedural texts : medical pratice guidelines. A linguistic study shows the relevancy and the limits of the structural clues to delimit the condition-action units, which form the basic semantic units for the segmentation task.
Dans cet article, nous etudions le role de la structure visuelle pour la segmentation de textes proceduraux. Nous nous focalisons sur un type de textes proceduraux particulier : les Guides de Bonnes Pratiques medicales. Une etude linguistique effectuee sur ce corpus montre la pertinence ainsi que les limites des indices visuels, pour delimiter des ensembles conditions-actions, qui forment des unites semantiques de base pour la segmentation. Cette etude a permis de definir une architecture modulaire qui exploite ces indices pour segmenter et structurer les textes.