SimText: a text mining framework for interactive analysis and visualization of similarities among biomedical entities
Journal
Bioinformatics
ISSN
1367-4803
1367-4811
Date Issued
2021
Author(s)
Marie Macnee
Sarah Schumacher-Bass
Jarrod Dalton
Costin Leu
Daniel Blankenberg
Dennis Lal
Type
Resource Types::text::journal::journal article
URL Institutional Repository
Abstract
<jats:title>Abstract</jats:title>
<jats:sec>
<jats:title>Summary</jats:title>
<jats:p>Literature exploration in PubMed on a large number of biomedical entities (e.g. genes, diseases or experiments) can be time-consuming and challenging, especially when assessing associations between entities. Here, we describe SimText, a user-friendly toolset that provides customizable and systematic workflows for the analysis of similarities among a set of entities based on text. SimText can be used for (i) text collection from PubMed and extraction of words with different text mining approaches, and (ii) interactive analysis and visualization of data using unsupervised learning techniques in an interactive app.</jats:p>
</jats:sec>
<jats:sec>
<jats:title>Availability and implementation</jats:title>
<jats:p>We developed SimText as an open-source R software and integrated it into Galaxy (https://usegalaxy.eu), an online data analysis platform with supporting self-learning training material available at https://training.galaxyproject.org. A command-line version of the toolset is available for download from GitHub (https://github.com/dlal-group/simtext) or as Docker image (https://hub.docker.com/r/dlalgroup/simtext/tags.).</jats:p>
</jats:sec>
<jats:sec>
<jats:title>Supplementary information</jats:title>
<jats:p>Supplementary data are available at Bioinformatics online.</jats:p>
</jats:sec>
<jats:sec>
<jats:title>Summary</jats:title>
<jats:p>Literature exploration in PubMed on a large number of biomedical entities (e.g. genes, diseases or experiments) can be time-consuming and challenging, especially when assessing associations between entities. Here, we describe SimText, a user-friendly toolset that provides customizable and systematic workflows for the analysis of similarities among a set of entities based on text. SimText can be used for (i) text collection from PubMed and extraction of words with different text mining approaches, and (ii) interactive analysis and visualization of data using unsupervised learning techniques in an interactive app.</jats:p>
</jats:sec>
<jats:sec>
<jats:title>Availability and implementation</jats:title>
<jats:p>We developed SimText as an open-source R software and integrated it into Galaxy (https://usegalaxy.eu), an online data analysis platform with supporting self-learning training material available at https://training.galaxyproject.org. A command-line version of the toolset is available for download from GitHub (https://github.com/dlal-group/simtext) or as Docker image (https://hub.docker.com/r/dlalgroup/simtext/tags.).</jats:p>
</jats:sec>
<jats:sec>
<jats:title>Supplementary information</jats:title>
<jats:p>Supplementary data are available at Bioinformatics online.</jats:p>
</jats:sec>
Cite this document
Macnee, M., Pérez-Palma, E., Schumacher-Bass, S., Dalton, J., Leu, C., Blankenberg, D., & Lal, D. (2021). SimText: A text mining framework for interactive analysis and visualization of similarities among biomedical entities. Bioinformatics, 37(22), 4285-4287. https://doi.org/10.1093/bioinformatics/btab365
Subjects
data analysis
;
data interpretation, statistical
;
data mining
;
pubmed
;
software
;
data analysis
;
data mining
;
medline
;
procedures
;
software
;
statistical analysis