CRIS
Permanent URI for this communityhttps://investigadores.udd.cl/handle/123456789/1
Browse
5 results
Search Results
Now showing 1 - 5 of 5
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Developing and Validating an Automatic Support System for Tumor Coding in Pathology Reports in Spanish(American Society of Clinical Oncology (ASCO), 2025-02) ;Fabián Villena ;Pablo Báez ;Sergio Peñafiel ;Matías RojasInti ParedesPurpose Pathology reports provide valuable information for cancer registries to understand, plan, and implement strategies to mitigate the impact of cancer. However, coding essential information from unstructured reports is performed by experts in a time-consuming manual process. We developed and validated a novel two-step automatic coding system that first recognizes tumor morphology and topography mentions from free text and then suggests codes from the International Classification of Diseases for Oncology (ICD-O) in Spanish. Materials and Methods We created an annotated corpus of tumor morphology and topography mentions consisting of 1,101 documents. We combined it with the CANTEMIST corpus (Cancer Text Mining Shared Task). Specifically, we implemented a named entity recognition (NER) model using the bidirectional long short-term memory network-conditional random field architecture enhanced with a stacked embedding layer. We applied transfer learning from state-of-the-art pretrained language models to obtain high-quality contextual representations, thus improving the detection of entities. The mentions found using this model were subsequently oded using a search engine tailored to the ICD-O codes. Results Our NER models achieved an F1 score of 0.86 and 0.90 for tumor morphology and topography, respectively. The overall performance of our automatic coding system achieved an accuracy at five suggestions of 0.72 and 0.65 for tumor morphology and topography, respectively. Conclusion These results demonstrate the feasibility of implementing natural language processing tools in the routine of a cancer center to extract and code valuable information from pathology reports. Our recommender system allows reliable and transparent coding at the moment of consultation. This publication shares the annotated corpus in Spanish, annotation guidelines, and source code to reproduce our experiments.Scopus© Citations 3 1 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, SimText: a text mining framework for interactive analysis and visualization of similarities among biomedical entities(2021) ;Marie Macnee; ;Sarah Schumacher-Bass ;Jarrod DaltonCostin Leu<jats:title>Abstract</jats:title> <jats:sec> <jats:title>Summary</jats:title> <jats:p>Literature exploration in PubMed on a large number of biomedical entities (e.g. genes, diseases or experiments) can be time-consuming and challenging, especially when assessing associations between entities. Here, we describe SimText, a user-friendly toolset that provides customizable and systematic workflows for the analysis of similarities among a set of entities based on text. SimText can be used for (i) text collection from PubMed and extraction of words with different text mining approaches, and (ii) interactive analysis and visualization of data using unsupervised learning techniques in an interactive app.</jats:p> </jats:sec> <jats:sec> <jats:title>Availability and implementation</jats:title> <jats:p>We developed SimText as an open-source R software and integrated it into Galaxy (https://usegalaxy.eu), an online data analysis platform with supporting self-learning training material available at https://training.galaxyproject.org. A command-line version of the toolset is available for download from GitHub (https://github.com/dlal-group/simtext) or as Docker image (https://hub.docker.com/r/dlalgroup/simtext/tags.).</jats:p> </jats:sec> <jats:sec> <jats:title>Supplementary information</jats:title> <jats:p>Supplementary data are available at Bioinformatics online.</jats:p> </jats:sec>Scopus© Citations 5 5 1 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Understanding water disputes in Chile with text and data mining tools(2019); ; ;Diego Rivera Salazar ;Douglas AitkenDaniel Brieba1Scopus© Citations 16 2 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, 1Scopus© Citations 26 3 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Scopus© Citations 15 1