CODER: Knowledge infused cross-lingual medical term embedding for term normalization

11/05/2020
by   Zheng Yuan, et al.
0

We propose a novel medical term embedding method named CODER, which stands for mediCal knOwledge embeDded tErm Representation. CODER is designed for medical term normalization by providing close vector representations for terms that represent the same or similar concepts with multi-language support. CODER is trained on top of BERT (Devlin et al., 2018) with the innovation that token vector aggregation is trained using relations from the UMLS Metathesaurus (Bodenreider, 2004), which is a comprehensive medical knowledge graph with multi-language support. Training with relations injects medical knowledge into term embeddings and aims to provide better normalization performances and potentially better machine learning features. We evaluated CODER in term normalization, semantic similarity, and relation classification benchmarks, which showed that CODER outperformed various state-of-the-art biomedical word embeddings, concept embeddings, and contextual embeddings.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
03/24/2020

Can Embeddings Adequately Represent Medical Terminology? New Large-Scale Medical Term Similarity Datasets Have the Answer!

A large number of embeddings trained on medical data have emerged, but i...
research
11/09/2022

Combining Contrastive Learning and Knowledge Graph Embeddings to develop medical word embeddings for the Italian language

Word embeddings play a significant role in today's Natural Language Proc...
research
02/08/2018

Biomedical term normalization of EHRs with UMLS

This paper presents a novel prototype for biomedical term normalization ...
research
09/02/2023

Knowledge Graph Embeddings for Multi-Lingual Structured Representations of Radiology Reports

The way we analyse clinical texts has undergone major changes over the l...
research
04/21/2019

Understanding Stability of Medical Concept Embeddings: Analysis and Prediction

In biomedical area, medical concepts linked to external knowledge bases ...
research
04/01/2022

Automatic Biomedical Term Clustering by Learning Fine-grained Term Representations

Term clustering is important in biomedical knowledge graph construction....
research
06/07/2020

Medical Concept Normalization in User Generated Texts by Learning Target Concept Embeddings

Medical concept normalization helps in discovering standard concepts in ...

Please sign up or login with your details

Forgot password? Click here to reset