liu.seSearch for publications in DiVA
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Eco-Friendly Witches: Improving Generalization Through Suspension of Disbelief for Zero-Shot Fact Extraction
Linköping University, Department of Computer and Information Science, Human-Centered Systems. Linköping University, Faculty of Science & Engineering.
Linköping University, Department of Computer and Information Science, Artificial Intelligence and Integrated Computer Systems. Linköping University, Faculty of Science & Engineering.
Linköping University, Department of Computer and Information Science, Human-Centered Systems. Linköping University, Faculty of Science & Engineering.ORCID iD: 0000-0003-0036-6662
2025 (English)In: PROCEEDINGS OF THE 13TH KNOWLEDGE CAPTURE CONFERENCE 2025, K-CAP 2025, Association for Computing Machinery (ACM), 2025, p. 44-51Conference paper, Published paper (Refereed)
Abstract [en]

As language models (LMs) increase in size and complexity, so too does their capacity to capture and recall knowledge from their pre-training data. This property is in some cases desirable, such as when applied to general-domain knowledge graph completion and relation extraction (RE) tasks. However, when extracting information about well-known entities from documents, it can be unclear whether LM-based systems perform well by actually solving the tasks, or instead by leveraging patterns about those entities learned from pre-training data. In realistic scenarios, documents may be domain-specific and unlike anything in the pre-training data, or may use well-known named entities in new ways. In the latter case, LMs need to disregard parts of their background knowledge to succeed, a capability we informally call suspension of disbelief. This paper uses a zero-shot RE setting to investigate whether masked LMs (MLMs) are capable of this. We devise a method called DocShRED to construct adversarial versions of well-known RE data sets like DocRED, by intentionally using named entities in ways that are inconsistent with common pre-training data. We find that BERT and RoBERTa exhibit near-random performance on DocShRED, even when using a zero-shot technique to incorporate the text as supporting information. We also find that both MLMs perform significantly better on DocShRED when entities' surface forms are withheld using a novel method called entity isolation, highlighting the impact of background knowledge on the task. Finally, when we apply entity isolation to the biomedical RE data set BioRED, BERT outperforms BioBERT and PubMedBERT without fine-tuning, suggesting an overall improvement in generalizability.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM), 2025. p. 44-51
Keywords [en]
Knowledge Graphs; Masked Language Models; Machine Reading; Entity Representation; Document-level Relation Extraction
National Category
Other Computer and Information Science
Identifiers
URN: urn:nbn:se:liu:diva-221321DOI: 10.1145/3731443.3771346ISI: 001665578100007Scopus ID: 2-s2.0-105024939839ISBN: 9798400718670 (print)OAI: oai:DiVA.org:liu-221321DiVA, id: diva2:2039060
Conference
13th International Conference on Knowledge Capture-K-CAP, Dayton, OH, dec 10-12, 2025
Note

Funding Agencies|Swedish National Graduate School in Computer Science (CUGS) - Excellence Center at Linkoping-Lund in Information Technology (ELLIIT)

Available from: 2026-02-16 Created: 2026-02-16 Last updated: 2026-08-18
In thesis
1. Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
Open this publication in new window or tab >>Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Modern knowledge-centric systems often use knowledge graphs (KGs) to represent information in a structured, queryable way, with the semantics underlying that structure described by an ontology. However, due to how rapidly knowledge can change in many real-world applications, handcrafting or even manually maintaining these KGs is not scalable. If the KGs are backed by a known data source, such as a collection of documents, then one solution is to partially automate the process, for example by automatically extracting facts from newly added documents. Ideally, such automation is accurate and faithful to the contents of each document. However, since automatic fact extraction from text often relies on language models, the results can be inaccurate for many reasons, such as underlying biases or a lack of pre-training data for a particular domain. This can be especially challenging in industry scenarios, where information from a variety of sources may need to be modeled, such as legal documents or scientific studies. This thesis aims to better understand the source of these inaccuracies and address them for encoder-only masked language models (MLMs), answering the question How can ontology-guided document modeling be performed using encoder-only MLMs?

The main contributions of this thesis are three-fold. First, it develops a framework for evaluating the entity representations produced by MLMs using a document-level relation extraction (DLRE) setting. In this setting, each document can be represented by a KG that describes its contents according to some ontology. Possible subject-relation-object triples from each document are verbalized into candidate statements, then scored by a given MLM and ranked, where that ranking is analyzed in various ways. By changing only the input representations of the subject and object entities of these statements and keeping all else constant, the experimental framework enables experiments that look at the effects of changes to those representations. Second, the thesis presents two zero-shot methods for improving those entity representations relative to document contents, both of which do not update the underlying MLM in any way. These representations were evaluated on three DLRE datasets based on Wikipedia articles, news articles, and biomedical paper abstracts. In one case, general-domain MLMs that used these improved representations even performed better on biomedical data than biomedical-specific MLMs that did not use them. Finally, the thesis develops a procedure for modifying general-domain DLRE datasets to be adversarial by intentionally using well-known entities in unusual contexts. The thesis found that MLMs overwhelmingly favored their background knowledge rather than the text of a document when performing statement ranking, so this adversarial data served to ensure that any improvements on the statement-ranking task were due to MLMs correctly using the documents’ text. These contributions are implemented in the publicly-available software package LINTEXT.

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2026. p. 85
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 2536
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-226764 (URN)10.3384/9789181186321 (DOI)9789181186314 (ISBN)9789181186321 (ISBN)
Public defence
2026-09-15, Ada Lovelace, B-building, Campus Valla, Linköping, 13:15 (English)
Opponent
Supervisors
Note

Funding: This work was funded in part by the Swedish National Graduate School in Computer Science (CUGS).

Available from: 2026-08-18 Created: 2026-08-18 Last updated: 2026-08-18Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Bueff, Andreas

Search in DiVA

By author/editor
Capshaw, RileyBueff, AndreasBlomqvist, Eva
By organisation
Human-Centered SystemsFaculty of Science & EngineeringArtificial Intelligence and Integrated Computer Systems
Other Computer and Information Science

Search outside of DiVA

GoogleGoogle Scholar

doi
isbn
urn-nbn

Altmetric score

doi
isbn
urn-nbn
Total: 72 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf