SynNER: Syntax-infused Named Entity Recognition in the Biomedical Domain
| UDC.coleccion | Investigación | |
| UDC.departamento | Ciencias da Computación e Tecnoloxías da Información | |
| UDC.grupoInv | Lingua e Sociedade da Información (LYS) | |
| UDC.institutoCentro | CITIC - Centro de Investigación de Tecnoloxías da Información e da Comunicación | |
| UDC.issue | 1 | |
| UDC.journalTitle | JAMIA Open | |
| UDC.volume | 9 | |
| dc.contributor.author | Imran, Muhammad | |
| dc.contributor.author | Zamaraeva, Olga | |
| dc.contributor.author | Gómez-Rodríguez, Carlos | |
| dc.date.accessioned | 2026-04-08T09:33:38Z | |
| dc.date.available | 2026-04-08T09:33:38Z | |
| dc.date.issued | 2026-02-21 | |
| dc.description | Financiado para publicación en acceso aberto: Universidade da Coruña/CISUG The code and other resources developed for this work are available in our GitHub repository at: https://github.com/chimran135/SynNER. The datasets used in this study are publicly available from third-party sources. The MTSamples and VAERS datasets can be downloaded from: https://github.com/BIDS-Xu-Lab/Clinical_Entity_Recognition_Using_GPT_models. The NCBI-Disease, BC2GM, and JNLPBA datasets can be downloaded from: https://github.com/cambridgeltl/MTL-Bioinformatics-2016. Supplementary material is available at JAMIA Open online | |
| dc.description.abstract | [Abstract]: Named Entity Recognition (NER) is a technology that helps computers automatically find and classify important terms in text, such as names of diseases, drugs, or medical procedures. This is especially valuable in the biomedical field, where researchers and clinicians need to process large volumes of text from scientific articles, clinical notes, or patient records. In this work, we present SynNER, a system that improves the accuracy of NER by teaching computers to pay attention not only to the words themselves, but also to syntax, ie, the internal structure of sentences. For example, recognizing how words are connected in a sentence (which parts of the sentence are subjects, objects or modifiers) can make it easier to correctly identify medical terms, even when they appear in complex contexts. We tested our method on five different collections of biomedical texts. The results showed that incorporating grammatical knowledge significantly boosted accuracy. Our SynNER system outperformed previous state-of-the-art methods on three of the five datasets. Our results show that using syntax can help researchers and healthcare professionals more reliably and quickly extract vital information from a vast amount of text, which could ultimately help improve biomedical research and clinical decision support tools. | |
| dc.description.sponsorship | We acknowledge the European Research Council (ERC), which has funded this research under the Horizon Europe research and innovation programme (SALSA, grant agreement No 101100615), SCANNER-UDC (PID2020-113230RB-C21) funded by MICIU/AEI/10.13039/501100011033, LATCHING (PID2023-147129OB-C21) funded by MICIU/AEI/10.13039/501100011033 and ERDF (EU), Ministry for Digital Transformation and Civil Service and “NextGenerationEU” PRTR under grant TSI-100925-2023-1, Xunta de Galicia (ED431C 2024/02), and Galician Research Center “CITIC”, funded by Xunta de Galicia through the collaboration agreement between the Consellería de Cultura, Educación, Formación Profesional e Universidades and the Galician universities for the reinforcement of the research centres of the Galician University System (CIGUS). Furthermore, this research was supported by the International, Interdisciplinary and Intersectoral Information and Communications Technology PhD programme (3-i ICT) granted to CITIC and supported by the European Union through the Horizon 2020 research and innovation programme under a Marie Skłodowska-Curie agreement (H2020-MSCA-COFUND), GA 101034261. Funding for open access charge: Universidade da Coruña/CISUG. | |
| dc.description.sponsorship | Xunta de Galicia; ED431C 2024/02 | |
| dc.identifier.citation | Muhammad Imran, Olga Zamaraeva, Carlos Gómez-Rodríguez, SynNER: syntax-infused named entity recognition in the biomedical domain, JAMIA Open, Volume 9, Issue 1, February 2026, ooaf149, https://doi.org/10.1093/jamiaopen/ooaf149 | |
| dc.identifier.doi | 10.1093/jamiaopen/ooaf149 | |
| dc.identifier.issn | 2574-2531 | |
| dc.identifier.uri | https://hdl.handle.net/2183/47894 | |
| dc.language.iso | eng | |
| dc.publisher | Oxford | |
| dc.relation.isbasedon | https://github.com/BIDS-Xu-Lab/Clinical_Entity_Recognition_Using_GPT_models | |
| dc.relation.isbasedon | https://github.com/cambridgeltl/MTL-Bioinformatics-2016 | |
| dc.relation.projectID | info:eu-repo/grantAgreement/EC/H2020/101034261 | |
| dc.relation.projectID | info:eu-repo/grantAgreement/EC/HE/101100615 | |
| dc.relation.projectID | info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2017-2020/PID2020-113230RB-C21/ES/MODELOS MULTITAREA DE ETIQUETADO SECUENCIAL PARA EL RECONOCIMIENTO DE ENTIDADES ENRIQUECIDO CON INFORMACION LINGUISTICA: SINTAXIS E INTEGRACION MULTITAREA (SCANNER-UDC) | |
| dc.relation.projectID | info:eu-repo/grantAgreement/AEI/Plan Estatal de Investigación Científica y Técnica y de Innovación 2021-2023/PID2023-147129OB-C21/ES/TECNOLOGÍAS DEL LENGUAJE DESDE UNA PERSPECTIVA VERDE (LATCHING): DOMINIOS CON ESCASOS RECURSOS | |
| dc.relation.projectID | info:eu-repo/grantAgreement/MTDPF//TSI-100925-2023-1/ES/CÁTEDRA UDC-INDITEX DE IA EN ALGORITMOS VERDES | |
| dc.relation.uri | https://doi.org/10.1093/jamiaopen/ooaf149 | |
| dc.rights | Attribution 4.0 International | en |
| dc.rights.accessRights | open access | |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | |
| dc.subject | Named entity recognition | |
| dc.subject | Dependency parsing | |
| dc.subject | Sequence labelling | |
| dc.title | SynNER: Syntax-infused Named Entity Recognition in the Biomedical Domain | |
| dc.type | journal article | |
| dc.type.hasVersion | VoR | |
| dspace.entity.type | Publication | |
| relation.isAuthorOfPublication | 6779b734-3d4b-4242-9bde-78e83eea84db | |
| relation.isAuthorOfPublication | e70a3969-39f6-4458-9339-3b71756fa56e | |
| relation.isAuthorOfPublication.latestForDiscovery | 6779b734-3d4b-4242-9bde-78e83eea84db |
Files
Original bundle
1 - 1 of 1
Loading...
- Name:
- GomezRodriguez_Carlos_2026_SynNER.pdf
- Size:
- 1.46 MB
- Format:
- Adobe Portable Document Format

