CRIS
Permanent URI for this communityhttps://investigadores.udd.cl/handle/123456789/1
Browse
4 results
Search Results
Now showing 1 - 4 of 4
- Some of the metrics are blocked by yourconsent settings
Item type:Publication, Developing and Validating an Automatic Support System for Tumor Coding in Pathology Reports in Spanish(American Society of Clinical Oncology (ASCO), 2025-02) ;Fabián Villena ;Pablo Báez ;Sergio Peñafiel ;Matías RojasInti ParedesPurpose Pathology reports provide valuable information for cancer registries to understand, plan, and implement strategies to mitigate the impact of cancer. However, coding essential information from unstructured reports is performed by experts in a time-consuming manual process. We developed and validated a novel two-step automatic coding system that first recognizes tumor morphology and topography mentions from free text and then suggests codes from the International Classification of Diseases for Oncology (ICD-O) in Spanish. Materials and Methods We created an annotated corpus of tumor morphology and topography mentions consisting of 1,101 documents. We combined it with the CANTEMIST corpus (Cancer Text Mining Shared Task). Specifically, we implemented a named entity recognition (NER) model using the bidirectional long short-term memory network-conditional random field architecture enhanced with a stacked embedding layer. We applied transfer learning from state-of-the-art pretrained language models to obtain high-quality contextual representations, thus improving the detection of entities. The mentions found using this model were subsequently oded using a search engine tailored to the ICD-O codes. Results Our NER models achieved an F1 score of 0.86 and 0.90 for tumor morphology and topography, respectively. The overall performance of our automatic coding system achieved an accuracy at five suggestions of 0.72 and 0.65 for tumor morphology and topography, respectively. Conclusion These results demonstrate the feasibility of implementing natural language processing tools in the routine of a cancer center to extract and code valuable information from pathology reports. Our recommender system allows reliable and transparent coding at the moment of consultation. This publication shares the annotated corpus in Spanish, annotation guidelines, and source code to reproduce our experiments.Scopus© Citations 3 1 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Large scale summarization using ensemble prompts and in context learning approaches(Springer Science and Business Media LLC, 2025-03-25) ;Andrés Leiva-Araos ;Bady Gana ;Héctor Allende-Cid ;José GarcíaManob Jyoti Saikia5 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Relevance of Machine Learning Techniques in Water Infrastructure Integrity and Quality: A Review Powered by Natural Language Processing(2023) ;José García ;LEIVA ARAOS, ANDRÉS ;Emerson Diaz-Saavedra ;Paola MoragaHernan PintoWater infrastructure integrity, quality, and distribution are fundamental for public health, environmental sustainability, economic development, and climate change resilience. Ensuring the robustness and quality of water infrastructure is pivotal for sectors like agriculture, industry, and energy production. Machine learning (ML) offers potential for bolstering water infrastructure integrity and quality by analyzing extensive data from sensors and other sources, optimizing treatment protocols, minimizing water losses, and improving distribution methods. This study delves into ML applications in water infrastructure integrity and quality by analyzing English-language articles from 2015 onward, compiling a total of 1087 articles. Initially, a natural language processing approach centered on topic modeling was adopted to classify salient topics. From each identified topic, key terms were extracted and utilized in a semi-automatic selection process, pinpointing the most relevant articles for further scrutiny, while unsupervised ML algorithms can assist in extracting themes from the documents, generating meaningful topics often requires intricate hyperparameter adjustments. Leveraging the Bidirectional Encoder Representations from Transformers (BERTopic) enhanced the study’s contextual comprehension in topic modeling. This semi-automatic methodology for bibliographic exploration begins with a broad topic categorization, advancing to an exhaustive analysis of each topic. The insights drawn underscore ML’s instrumental role in enhancing water infrastructure’s integrity and quality, suggesting promising future research directions. Specifically, the study has identified four key areas where ML has been applied to water management: (1) advancements in the detection of water contaminants and soil erosion; (2) forecasting of water levels; (3) advanced techniques for leak detection in water networks; and (4) evaluation of water quality and potability. These findings underscore the transformative impact of ML on water infrastructure and suggest promising paths for continued investigation.Scopus© Citations 12 1 - Some of the metrics are blocked by yourconsent settings
Item type:Publication, Textual inference for eligibility criteria resolution in clinical trials(2015) ;Chaitanya Shivade ;Courtney Hebert ;Marcelo Lopetegui ;Marie-Catherine de MarneffeEric Fosler-Lussier8Scopus© Citations 31