SynNER: Syntax-infused Named Entity Recognition in the Biomedical Domain

UDC.coleccionInvestigación
UDC.departamentoCiencias da Computación e Tecnoloxías da Información
UDC.grupoInvLingua e Sociedade da Información (LYS)
UDC.institutoCentroCITIC - Centro de Investigación de Tecnoloxías da Información e da Comunicación
UDC.issue1
UDC.journalTitleJAMIA Open
UDC.volume9
dc.contributor.authorImran, Muhammad
dc.contributor.authorZamaraeva, Olga
dc.contributor.authorGómez-Rodríguez, Carlos
dc.date.accessioned2026-04-08T09:33:38Z
dc.date.available2026-04-08T09:33:38Z
dc.date.issued2026-02-21
dc.descriptionFinanciado para publicación en acceso aberto: Universidade da Coruña/CISUG The code and other resources developed for this work are available in our GitHub repository at: https://github.com/chimran135/SynNER. The datasets used in this study are publicly available from third-party sources. The MTSamples and VAERS datasets can be downloaded from: https://github.com/BIDS-Xu-Lab/Clinical_Entity_Recognition_Using_GPT_models. The NCBI-Disease, BC2GM, and JNLPBA datasets can be downloaded from: https://github.com/cambridgeltl/MTL-Bioinformatics-2016. Supplementary material is available at JAMIA Open online
dc.description.abstract[Abstract]: Named Entity Recognition (NER) is a technology that helps computers automatically find and classify important terms in text, such as names of diseases, drugs, or medical procedures. This is especially valuable in the biomedical field, where researchers and clinicians need to process large volumes of text from scientific articles, clinical notes, or patient records. In this work, we present SynNER, a system that improves the accuracy of NER by teaching computers to pay attention not only to the words themselves, but also to syntax, ie, the internal structure of sentences. For example, recognizing how words are connected in a sentence (which parts of the sentence are subjects, objects or modifiers) can make it easier to correctly identify medical terms, even when they appear in complex contexts. We tested our method on five different collections of biomedical texts. The results showed that incorporating grammatical knowledge significantly boosted accuracy. Our SynNER system outperformed previous state-of-the-art methods on three of the five datasets. Our results show that using syntax can help researchers and healthcare professionals more reliably and quickly extract vital information from a vast amount of text, which could ultimately help improve biomedical research and clinical decision support tools.
dc.description.sponsorshipWe acknowledge the European Research Council (ERC), which has funded this research under the Horizon Europe research and innovation programme (SALSA, grant agreement No 101100615), SCANNER-UDC (PID2020-113230RB-C21) funded by MICIU/AEI/10.13039/501100011033, LATCHING (PID2023-147129OB-C21) funded by MICIU/AEI/10.13039/501100011033 and ERDF (EU), Ministry for Digital Transformation and Civil Service and “NextGenerationEU” PRTR under grant TSI-100925-2023-1, Xunta de Galicia (ED431C 2024/02), and Galician Research Center “CITIC”, funded by Xunta de Galicia through the collaboration agreement between the Consellería de Cultura, Educación, Formación Profesional e Universidades and the Galician universities for the reinforcement of the research centres of the Galician University System (CIGUS). Furthermore, this research was supported by the International, Interdisciplinary and Intersectoral Information and Communications Technology PhD programme (3-i ICT) granted to CITIC and supported by the European Union through the Horizon 2020 research and innovation programme under a Marie Skłodowska-Curie agreement (H2020-MSCA-COFUND), GA 101034261. Funding for open access charge: Universidade da Coruña/CISUG.
dc.description.sponsorshipXunta de Galicia; ED431C 2024/02
dc.identifier.citationMuhammad Imran, Olga Zamaraeva, Carlos Gómez-Rodríguez, SynNER: syntax-infused named entity recognition in the biomedical domain, JAMIA Open, Volume 9, Issue 1, February 2026, ooaf149, https://doi.org/10.1093/jamiaopen/ooaf149
dc.identifier.doi10.1093/jamiaopen/ooaf149
dc.identifier.issn2574-2531
dc.identifier.urihttps://hdl.handle.net/2183/47894
dc.language.isoeng
dc.publisherOxford
dc.relation.isbasedonhttps://github.com/BIDS-Xu-Lab/Clinical_Entity_Recognition_Using_GPT_models
dc.relation.isbasedonhttps://github.com/cambridgeltl/MTL-Bioinformatics-2016
dc.relation.projectIDinfo:eu-repo/grantAgreement/EC/H2020/101034261
dc.relation.projectIDinfo:eu-repo/grantAgreement/EC/HE/101100615
dc.relation.projectIDinfo:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020/PID2020-113230RB-C21/ES/MODELOS MULTITAREA DE ETIQUETADO SECUENCIAL PARA EL RECONOCIMIENTO DE ENTIDADES ENRIQUECIDO CON INFORMACION LINGUISTICA: SINTAXIS E INTEGRACION MULTITAREA (SCANNER-UDC)
dc.relation.projectIDinfo:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/PID2023-147129OB-C21/ES/TECNOLOGÍAS DEL LENGUAJE DESDE UNA PERSPECTIVA VERDE (LATCHING): DOMINIOS CON ESCASOS RECURSOS
dc.relation.projectIDinfo:eu-repo/grantAgreement/MTDPF//TSI-100925-2023-1/ES/CÁTEDRA UDC-INDITEX DE IA EN ALGORITMOS VERDES
dc.relation.urihttps://doi.org/10.1093/jamiaopen/ooaf149
dc.rightsAttribution 4.0 Internationalen
dc.rights.accessRightsopen access
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.subjectNamed entity recognition
dc.subjectDependency parsing
dc.subjectSequence labelling
dc.titleSynNER: Syntax-infused Named Entity Recognition in the Biomedical Domain
dc.typejournal article
dc.type.hasVersionVoR
dspace.entity.typePublication
relation.isAuthorOfPublication6779b734-3d4b-4242-9bde-78e83eea84db
relation.isAuthorOfPublicatione70a3969-39f6-4458-9339-3b71756fa56e
relation.isAuthorOfPublication.latestForDiscovery6779b734-3d4b-4242-9bde-78e83eea84db

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
GomezRodriguez_Carlos_2026_SynNER.pdf
Size:
1.46 MB
Format:
Adobe Portable Document Format