Enhancing Reliability in AI-Generated Academic References through Automated Verification.
2026 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE credits
Student thesis
Abstract [en]
Academic integrity depends on the accuracy of scholarly references. When referencesare fabricated, the credibility of a paper cannot be established, and the chain of knowledgethat connects new work to existing research breaks down. This creates a growing crisis oftrust, particularly as it becomes harder to distinguish carefully researched bibliographiesfrom those that were generated automatically.Students are increasingly using AI writing tools to assist with editing, cleaning updrafts, and finishing background sections. During this process, it can easily happen thatthe AI hallucinates inventing references that look real but point to publications that do notexist. In other cases, students rely on AI entirely for background research. Given the largenumber of references such tools can generate in a short time, it becomes practically impossible for supervisors and reviewers to verify them manually. Each reference must includethe author names, title, year, publication venue, and Digital Object Identifier (DOI), and allof these details must be correct. A single fabricated or corrupted entry can pass unnoticedthrough standard checks yet undermine the reliability of the entire work.We present BIBVERIFY, a fully automated agentic system that addresses this challenge.BIBVERIFY reads reference files in BibTeX format, verifies each entry against authoritativedatabases through the Model Context Protocol (MCP), and classifies them as valid, partiallyvalid, or invalid. It primarily uses the DBLP Computer Science Bibliography, with GoogleScholar as a fallback for references not covered by DBLP. For partially valid entries, it applies conservative, evidence-based corrections without any human involvement. Unlike aplain MCP agent that simply retrieves database records, BIBVERIFY combines structuredretrieval with LLM-based reasoning to detect subtle field-level errors such as misspelledauthor names, wrong years, or invalid DOIs that a lookup alone would miss.To assess the effectiveness of this approach, performance is evaluated across classification accuracy, issue detection, correction success rate, field-level accuracy for attributessuch as author names, titles, and DOIs, and overall runtime. Three prompting strategies zero-shot, retrieval-augmented generation (RAG), and chain-of-thought, are comparedacross different language models.
Place, publisher, year, edition, pages
2026. , p. 41
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:liu:diva-227194ISRN: LIU-IDA/LITH-EX-A--26/053--SEOAI: oai:DiVA.org:liu-227194DiVA, id: diva2:2096817
Presentation
2026-06-12, IDA Alan Turing 40, Linköping, 11:44 (English)
Supervisors
Examiners
2026-08-312026-08-312026-08-31Bibliographically approved