liu.seSearch for publications in DiVA
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
LINTEXT: A Visual Tool for Exploring and Modeling Knowledge in Text Documents
Linköping University, Department of Computer and Information Science, Human-Centered Systems. Linköping University, Faculty of Science & Engineering.ORCID iD: 0009-0009-3977-2706
Linköping University, Faculty of Science & Engineering. Linköping University, Department of Computer and Information Science, Human-Centered Systems.ORCID iD: 0000-0003-0036-6662
2024 (English)In: Joint Proceedings of Posters, Demos, Workshops, and Tutorials of the 24th International Conference on Knowledge Engineering and Knowledge Management (EKAW-PDWT 2024)co-located with 24th International Conference on Knowledge Engineering and Knowledge Management (EKAW 2024) / [ed] Carlos Badenes-Olmedo, Inna Novalija, Enrico Daga, Lise Stork, Reshmi Gopalakrishna Pillai, Laurence Dierickx, Benno Kruit, Victoria Degeler, João Moreira, Bohui Zhang, Reham Alharbi, Yuan He, Arianna Graciotti, Alba Morales Tirado, Valentina Presutti, Enrico Motta, CEUR-WS , 2024, Vol. 3967Conference paper, Published paper (Refereed)
Abstract [en]

A large part of knowledge is commonly encoded into text documents. While extracting this information into a Knowledge Graph (KG) is a common approach, it suffers from challenges when texts are added, removed, or changed, or when the schema of the intended KG changes. Instead we advocate an approach where text and models evolve together in an interactive manner. We present LINTEXT, a system accompanying a published method which allows users to jointly explore and model the information held within text documents. The modeling is accomplished by specifying fill-in-the-blank prompts along with some metadata which are then recorded as specifications for simple relations that can be used to generate an ontology. The exploration aspect is accomplished by having the system complete each prompt with entities identified from the text and presenting the completions as a ranked list to the user, allowing users to verify the quality of the extracted triples. By elevating the development of the ontology to a visual and interactive level, it has an immediate text connection and users can be more certain that the documents they wish to model contain the information they wish to extract or query. Additionally, our system is designed to support the development of relation extraction (RE) pipelines underlying the document analysis, with a particular focus on supporting methods for improving vector representations of the extracted entities. To this end, users can choose to analyze documents from pre-annotated RE data sets to understand how changes in different elements of the pipeline affect the results.

Place, publisher, year, edition, pages
CEUR-WS , 2024. Vol. 3967
Series
CEUR Workshop Proceedings, ISSN 1613-0073
Keywords [en]
Knowledge Graphs, Masked Language Models, Machine Reading, Entity Embedding, Document-level Relation Extraction, Interactive Knowledge Modeling
National Category
Computer Sciences Natural Language Processing
Identifiers
URN: urn:nbn:se:liu:diva-226750Scopus ID: 2-s2.0-105006925553OAI: oai:DiVA.org:liu-226750DiVA, id: diva2:2092705
Conference
Posters, Demos, Workshops, and Tutorials of the 24th International Conference on Knowledge Engineering and Knowledge Management (EKAW-PDWT 2024) co-located with 24th International Conference on Knowledge Engineering and Knowledge Management (EKAW 2024), Amsterdam, Netherlands, November 26-28, 2024.
Available from: 2026-08-17 Created: 2026-08-17 Last updated: 2026-08-18Bibliographically approved
In thesis
1. Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
Open this publication in new window or tab >>Composing Meaning from Text: Improving Entity Representations for Ontology-based Document Modeling using Masked Language Models
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Modern knowledge-centric systems often use knowledge graphs (KGs) to represent information in a structured, queryable way, with the semantics underlying that structure described by an ontology. However, due to how rapidly knowledge can change in many real-world applications, handcrafting or even manually maintaining these KGs is not scalable. If the KGs are backed by a known data source, such as a collection of documents, then one solution is to partially automate the process, for example by automatically extracting facts from newly added documents. Ideally, such automation is accurate and faithful to the contents of each document. However, since automatic fact extraction from text often relies on language models, the results can be inaccurate for many reasons, such as underlying biases or a lack of pre-training data for a particular domain. This can be especially challenging in industry scenarios, where information from a variety of sources may need to be modeled, such as legal documents or scientific studies. This thesis aims to better understand the source of these inaccuracies and address them for encoder-only masked language models (MLMs), answering the question How can ontology-guided document modeling be performed using encoder-only MLMs?

The main contributions of this thesis are three-fold. First, it develops a framework for evaluating the entity representations produced by MLMs using a document-level relation extraction (DLRE) setting. In this setting, each document can be represented by a KG that describes its contents according to some ontology. Possible subject-relation-object triples from each document are verbalized into candidate statements, then scored by a given MLM and ranked, where that ranking is analyzed in various ways. By changing only the input representations of the subject and object entities of these statements and keeping all else constant, the experimental framework enables experiments that look at the effects of changes to those representations. Second, the thesis presents two zero-shot methods for improving those entity representations relative to document contents, both of which do not update the underlying MLM in any way. These representations were evaluated on three DLRE datasets based on Wikipedia articles, news articles, and biomedical paper abstracts. In one case, general-domain MLMs that used these improved representations even performed better on biomedical data than biomedical-specific MLMs that did not use them. Finally, the thesis develops a procedure for modifying general-domain DLRE datasets to be adversarial by intentionally using well-known entities in unusual contexts. The thesis found that MLMs overwhelmingly favored their background knowledge rather than the text of a document when performing statement ranking, so this adversarial data served to ensure that any improvements on the statement-ranking task were due to MLMs correctly using the documents’ text. These contributions are implemented in the publicly-available software package LINTEXT.

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2026. p. 85
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 2536
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-226764 (URN)10.3384/9789181186321 (DOI)9789181186314 (ISBN)9789181186321 (ISBN)
Public defence
2026-09-15, Ada Lovelace, B-building, Campus Valla, Linköping, 13:15 (English)
Opponent
Supervisors
Note

Funding: This work was funded in part by the Swedish National Graduate School in Computer Science (CUGS).

Available from: 2026-08-18 Created: 2026-08-18 Last updated: 2026-08-18Bibliographically approved

Open Access in DiVA

fulltext(662 kB)11 downloads
File information
File name FULLTEXT01.pdfFile size 662 kBChecksum SHA-512
40b30bf2686af40d50b1d70fc9fd4ea8697b25e8f6d46c676eaff2e658d74787f09d7fe4a6d299378d1af03138e3636715449e8e6859de03ed7e4f6650a987eb
Type fulltextMimetype application/pdf

Other links

ScopusFulltext

Authority records

Capshaw, RileyBlomqvist, Eva

Search in DiVA

By author/editor
Capshaw, RileyBlomqvist, Eva
By organisation
Human-Centered SystemsFaculty of Science & Engineering
Computer SciencesNatural Language Processing

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 3379 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf