liu.seSearch for publications in DiVA
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Contextualizing Entity Representations for Zero-Shot Relation Extraction with Masked Language Models
Linköping University, Department of Computer and Information Science, Human-Centered Systems. Linköping University, Faculty of Science & Engineering.
Linköping University, Department of Computer and Information Science, Human-Centered Systems. Linköping University, Faculty of Science & Engineering.ORCID iD: 0000-0003-0036-6662
2025 (English)In: KNOWLEDGE ENGINEERING AND KNOWLEDGE MANAGEMENT, EKAW 2024, SPRINGER INTERNATIONAL PUBLISHING AG , 2025, Vol. 15370, p. 399-415Conference paper, Published paper (Refereed)
Abstract [en]

Knowledge graphs (KGs) and their related ontologies constitute a key component in modern knowledge-based systems. However, hand-crafting these is not scalable, particularly due to the rate at which knowledge changes in many real-world applications. Partially automating the process of extracting and even modelling knowledge has therefore been a subject of research for many years. Nevertheless, accurate and reliable KG construction from natural language documents still remains a difficult task with many challenges, even in light of the impressive recent advances in language modelling. This paper focuses on one of those challenges, namely the extraction of accurate entity representations from text documents in order to facilitate relation extraction (RE). We present a novel method for generating document-contextualized input representations for entities using a masked language model (MLM) without the need for any sort of fine-tuning. These representations are then used as inputs to the same MLM that generated them, alleviating the need to include entire documents when prompting. Our results show that these representations 1) improve the ability of the MLMs BERT and RoBERTa to identify statements that represent correct relations between two entities; and 2) allow BERT to perform on par with the fine-tuned MLMs BioBERT and PubMedBERT.

Place, publisher, year, edition, pages
SPRINGER INTERNATIONAL PUBLISHING AG , 2025. Vol. 15370, p. 399-415
Series
Lecture Notes in Artificial Intelligence, ISSN 2945-9133
Keywords [en]
Knowledge Graphs; Masked Language Models; Machine Reading; Entity Embedding; Document-level Relation Extraction
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:liu:diva-217990DOI: 10.1007/978-3-031-77792-9_24ISI: 001542675000024Scopus ID: 2-s2.0-85210848015ISBN: 9783031777912 (print)ISBN: 9783031777929 (electronic)OAI: oai:DiVA.org:liu-217990DiVA, id: diva2:2001458
Conference
24th International Conference on Knowledge Engineering and Knowledge Management-EKAW, Amsterdam, NETHERLANDS, nov 26-28, 2024
Note

Funding Agencies|Swedish National Graduate School in Computer Science (CUGS) - Excellence Center at Linkoping-Lund in Information Technology (ELLIIT)

Available from: 2025-09-26 Created: 2025-09-26 Last updated: 2026-08-18
In thesis
1. Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
Open this publication in new window or tab >>Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Modern knowledge-centric systems often use knowledge graphs (KGs) to represent information in a structured, queryable way, with the semantics underlying that structure described by an ontology. However, due to how rapidly knowledge can change in many real-world applications, handcrafting or even manually maintaining these KGs is not scalable. If the KGs are backed by a known data source, such as a collection of documents, then one solution is to partially automate the process, for example by automatically extracting facts from newly added documents. Ideally, such automation is accurate and faithful to the contents of each document. However, since automatic fact extraction from text often relies on language models, the results can be inaccurate for many reasons, such as underlying biases or a lack of pre-training data for a particular domain. This can be especially challenging in industry scenarios, where information from a variety of sources may need to be modeled, such as legal documents or scientific studies. This thesis aims to better understand the source of these inaccuracies and address them for encoder-only masked language models (MLMs), answering the question How can ontology-guided document modeling be performed using encoder-only MLMs?

The main contributions of this thesis are three-fold. First, it develops a framework for evaluating the entity representations produced by MLMs using a document-level relation extraction (DLRE) setting. In this setting, each document can be represented by a KG that describes its contents according to some ontology. Possible subject-relation-object triples from each document are verbalized into candidate statements, then scored by a given MLM and ranked, where that ranking is analyzed in various ways. By changing only the input representations of the subject and object entities of these statements and keeping all else constant, the experimental framework enables experiments that look at the effects of changes to those representations. Second, the thesis presents two zero-shot methods for improving those entity representations relative to document contents, both of which do not update the underlying MLM in any way. These representations were evaluated on three DLRE datasets based on Wikipedia articles, news articles, and biomedical paper abstracts. In one case, general-domain MLMs that used these improved representations even performed better on biomedical data than biomedical-specific MLMs that did not use them. Finally, the thesis develops a procedure for modifying general-domain DLRE datasets to be adversarial by intentionally using well-known entities in unusual contexts. The thesis found that MLMs overwhelmingly favored their background knowledge rather than the text of a document when performing statement ranking, so this adversarial data served to ensure that any improvements on the statement-ranking task were due to MLMs correctly using the documents’ text. These contributions are implemented in the publicly-available software package LINTEXT.

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2026. p. 85
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 2536
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-226764 (URN)10.3384/9789181186321 (DOI)9789181186314 (ISBN)9789181186321 (ISBN)
Public defence
2026-09-15, Ada Lovelace, B-building, Campus Valla, Linköping, 13:15 (English)
Opponent
Supervisors
Note

Funding: This work was funded in part by the Swedish National Graduate School in Computer Science (CUGS).

Available from: 2026-08-18 Created: 2026-08-18 Last updated: 2026-08-18Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Search in DiVA

By author/editor
Capshaw, RileyBlomqvist, Eva
By organisation
Human-Centered SystemsFaculty of Science & Engineering
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 53 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf