Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]
Modern knowledge-centric systems often use knowledge graphs (KGs) to represent information in a structured, queryable way, with the semantics underlying that structure described by an ontology. However, due to how rapidly knowledge can change in many real-world applications, handcrafting or even manually maintaining these KGs is not scalable. If the KGs are backed by a known data source, such as a collection of documents, then one solution is to partially automate the process, for example by automatically extracting facts from newly added documents. Ideally, such automation is accurate and faithful to the contents of each document. However, since automatic fact extraction from text often relies on language models, the results can be inaccurate for many reasons, such as underlying biases or a lack of pre-training data for a particular domain. This can be especially challenging in industry scenarios, where information from a variety of sources may need to be modeled, such as legal documents or scientific studies. This thesis aims to better understand the source of these inaccuracies and address them for encoder-only masked language models (MLMs), answering the question How can ontology-guided document modeling be performed using encoder-only MLMs?
The main contributions of this thesis are three-fold. First, it develops a framework for evaluating the entity representations produced by MLMs using a document-level relation extraction (DLRE) setting. In this setting, each document can be represented by a KG that describes its contents according to some ontology. Possible subject-relation-object triples from each document are verbalized into candidate statements, then scored by a given MLM and ranked, where that ranking is analyzed in various ways. By changing only the input representations of the subject and object entities of these statements and keeping all else constant, the experimental framework enables experiments that look at the effects of changes to those representations. Second, the thesis presents two zero-shot methods for improving those entity representations relative to document contents, both of which do not update the underlying MLM in any way. These representations were evaluated on three DLRE datasets based on Wikipedia articles, news articles, and biomedical paper abstracts. In one case, general-domain MLMs that used these improved representations even performed better on biomedical data than biomedical-specific MLMs that did not use them. Finally, the thesis develops a procedure for modifying general-domain DLRE datasets to be adversarial by intentionally using well-known entities in unusual contexts. The thesis found that MLMs overwhelmingly favored their background knowledge rather than the text of a document when performing statement ranking, so this adversarial data served to ensure that any improvements on the statement-ranking task were due to MLMs correctly using the documents’ text. These contributions are implemented in the publicly-available software package LINTEXT.
Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2026. , p. 85
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 2536
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:liu:diva-226764DOI: 10.3384/9789181186321ISBN: 9789181186314 (print)ISBN: 9789181186321 (electronic)OAI: oai:DiVA.org:liu-226764DiVA, id: diva2:2093014
Public defence
2026-09-15, Ada Lovelace, B-building, Campus Valla, Linköping, 13:15 (English)
Opponent
Supervisors
Note
Funding: This work was funded in part by the Swedish National Graduate School in Computer Science (CUGS).
2026-08-182026-08-182026-08-18Bibliographically approved
List of papers