liu.seSearch for publications in DiVA
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Towards Tailored Knowledge Base Modeling using Masked Language Models
Linköping University, Department of Computer and Information Science, Human-Centered Systems. Linköping University, Faculty of Science & Engineering.ORCID iD: 0009-0009-3977-2706
Linköping University, Faculty of Science & Engineering. Linköping University, Department of Computer and Information Science, Human-Centered Systems.ORCID iD: 0000-0003-0036-6662
2023 (English)In: Joint Proceedings of the Second International Workshop on Knowledge Graph Generation From Text and the First International BiKE Challengeco-located with 20th Extended Semantic Conference (ESWC 2023) / [ed] Sanju Tiwari, Nandana Mihindukulasooriya, Francesco Osborne, Dimitris Kontokostas, Jennifer D’Souza, Mayank Kejriwal, Edgard Marx, CEUR-WS , 2023, Vol. 3447, p. 118-131Conference paper, Published paper (Refereed)
Abstract [en]

We propose a methodology for leveraging aspects of ontology design principles to guide the use of a masked language model (MLM) as a query engine over raw text documents. By using targeted fill-in-the blank-style prompts to define relations, we show how a domain expert could use BERT, a well-known MLM, to extract triples from unseen documents without any fine-tuning. We evaluate our proposed methodology using a modified document-level relation extraction task, highlighting early successes but also numerous areas that need improvement. Despite these shortcomings, we then discuss why we are still hopeful that this paves the way toward flexible text-based query engines which use collections of unstructured documents.

Place, publisher, year, edition, pages
CEUR-WS , 2023. Vol. 3447, p. 118-131
Series
CEUR Workshop Proceedings, ISSN 1613-0073 ; 3447
Keywords [en]
Knowledge Graphs, Masked Language Models, Ontologies, Document-level Relation Extraction
National Category
Natural Language Processing
Identifiers
URN: urn:nbn:se:liu:diva-226749Scopus ID: 2-s2.0-85169151570OAI: oai:DiVA.org:liu-226749DiVA, id: diva2:2092679
Conference
TEXT2KG 2023 & BiKE 2023:Second International Workshop on Knowledge Graph Generation From Text and the First International BiKE Challenge, co-located with 20th Extended Semantic Conference (ESWC 2023), Hersonissos, Greece, May 29th, 2023.
Available from: 2026-08-17 Created: 2026-08-17 Last updated: 2026-08-18Bibliographically approved
In thesis
1. Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
Open this publication in new window or tab >>Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Modern knowledge-centric systems often use knowledge graphs (KGs) to represent information in a structured, queryable way, with the semantics underlying that structure described by an ontology. However, due to how rapidly knowledge can change in many real-world applications, handcrafting or even manually maintaining these KGs is not scalable. If the KGs are backed by a known data source, such as a collection of documents, then one solution is to partially automate the process, for example by automatically extracting facts from newly added documents. Ideally, such automation is accurate and faithful to the contents of each document. However, since automatic fact extraction from text often relies on language models, the results can be inaccurate for many reasons, such as underlying biases or a lack of pre-training data for a particular domain. This can be especially challenging in industry scenarios, where information from a variety of sources may need to be modeled, such as legal documents or scientific studies. This thesis aims to better understand the source of these inaccuracies and address them for encoder-only masked language models (MLMs), answering the question How can ontology-guided document modeling be performed using encoder-only MLMs?

The main contributions of this thesis are three-fold. First, it develops a framework for evaluating the entity representations produced by MLMs using a document-level relation extraction (DLRE) setting. In this setting, each document can be represented by a KG that describes its contents according to some ontology. Possible subject-relation-object triples from each document are verbalized into candidate statements, then scored by a given MLM and ranked, where that ranking is analyzed in various ways. By changing only the input representations of the subject and object entities of these statements and keeping all else constant, the experimental framework enables experiments that look at the effects of changes to those representations. Second, the thesis presents two zero-shot methods for improving those entity representations relative to document contents, both of which do not update the underlying MLM in any way. These representations were evaluated on three DLRE datasets based on Wikipedia articles, news articles, and biomedical paper abstracts. In one case, general-domain MLMs that used these improved representations even performed better on biomedical data than biomedical-specific MLMs that did not use them. Finally, the thesis develops a procedure for modifying general-domain DLRE datasets to be adversarial by intentionally using well-known entities in unusual contexts. The thesis found that MLMs overwhelmingly favored their background knowledge rather than the text of a document when performing statement ranking, so this adversarial data served to ensure that any improvements on the statement-ranking task were due to MLMs correctly using the documents’ text. These contributions are implemented in the publicly-available software package LINTEXT.

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2026. p. 85
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 2536
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-226764 (URN)10.3384/9789181186321 (DOI)9789181186314 (ISBN)9789181186321 (ISBN)
Public defence
2026-09-15, Ada Lovelace, B-building, Campus Valla, Linköping, 13:15 (English)
Opponent
Supervisors
Note

Funding: This work was funded in part by the Swedish National Graduate School in Computer Science (CUGS).

Available from: 2026-08-18 Created: 2026-08-18 Last updated: 2026-08-18Bibliographically approved

Open Access in DiVA

fulltext(404 kB)13 downloads
File information
File name FULLTEXT01.pdfFile size 404 kBChecksum SHA-512
2969865683ffa52861ea653194d280f14990f510da49e6c5c4990cf0a18849f1ebc8639ab530c9b696deef3067ca14663aeb26464fc817f81e707752ce4db08b
Type fulltextMimetype application/pdf

Other links

ScopusFulltext

Authority records

Capshaw, RileyBlomqvist, Eva

Search in DiVA

By author/editor
Capshaw, RileyBlomqvist, Eva
By organisation
Human-Centered SystemsFaculty of Science & Engineering
Natural Language Processing

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 3256 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf