liu.seSök publikationer i DiVA
Ändra sökning
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Adaptive Semi-structured Information Extraction
Linköpings universitet, Institutionen för datavetenskap, KPLAB - Laboratoriet för kunskapsbearbetning. Linköpings universitet, Tekniska högskolan.
2003 (Engelska)Licentiatavhandling, monografi (Övrigt vetenskapligt)
Abstract [en]

The number of domains and tasks where information extraction tools can be used needs to be increased. One way to reach this goal is to construct user-driven information extraction systems where novice users are able to adapt them to new domains and tasks. To accomplish this goal, the systems need to become more intelligent and able to learn to extract information without need of expert skills or time-consuming work from the user.

The type of information extraction system that is in focus for this thesis is semistructural information extraction. The term semi-structural refers to documents that not only contain natural language text but also additional structural information. The typical application is information extraction from World Wide Web hypertext documents. By making effective use of not only the link structure but also the structural information within each such document, user-driven extraction systems with high performance can be built.

The extraction process contains several steps where different types of techniques are used. Examples of such types of techniques are those that take advantage of structural, pure syntactic, linguistic, and semantic information. The first step that is in focus for this thesis is the navigation step that takes advantage of the structural information. It is only one part of a complete extraction system, but it is an important part. The use of reinforcement learning algorithms for the navigation step can make the adaptation of the system to new tasks and domains more user-driven. The advantage of using reinforcement learning techniques is that the extraction agent can efficiently learn from its own experience without need for intensive user interactions.

An agent-oriented system was designed to evaluate the approach suggested in this thesis. Initial experiments showed that the training of the navigation step and the approach of the system was promising. However, additional components need to be included in the system before it becomes a fully-fledged user-driven system.

Ort, förlag, år, upplaga, sidor
Institutionen för datavetenskap , 2003. , s. 85
Serie
Linköping Studies in Science and Technology. Thesis, ISSN 0280-7971 ; 1000
Nyckelord [en]
Information extraction, Artificial intelligence, Semi-structured data, Reinforced learning, Knowledge management
Nationell ämneskategori
Datavetenskap (datalogi)
Identifikatorer
URN: urn:nbn:se:liu:diva-5688Lokalt ID: LiU-Tek-Lic-2002:73ISBN: 9173735892 (tryckt)OAI: oai:DiVA.org:liu-5688DiVA, id: diva2:21449
Presentation
2002-12-15, 00:00 (Engelska)
Handledare
Anmärkning

Report code: LiU-Tek-Lic-2002:73.

Tillgänglig från: 2003-01-30 Skapad: 2003-01-30 Senast uppdaterad: 2023-01-25Bibliografiskt granskad

Open Access i DiVA

fulltext(432 kB)1438 nedladdningar
Filinformation
Filnamn FULLTEXT01.pdfFilstorlek 432 kBChecksumma SHA-1
e5702c06b15b52e7b4564f08fefaa8cf1e2c3fd4390956952340eec4d949469ef3da7d1f
Typ fulltextMimetyp application/pdf

Person

Arpteg, Anders

Sök vidare i DiVA

Av författaren/redaktören
Arpteg, Anders
Av organisationen
KPLAB - Laboratoriet för kunskapsbearbetningTekniska högskolan
Datavetenskap (datalogi)

Sök vidare utanför DiVA

GoogleGoogle Scholar
Totalt: 1444 nedladdningar
Antalet nedladdningar är summan av nedladdningar för alla fulltexter. Det kan inkludera t.ex tidigare versioner som nu inte längre är tillgängliga.

isbn
urn-nbn

Altmetricpoäng

isbn
urn-nbn
Totalt: 1401 träffar
RefereraExporteraLänk till posten
Permanent länk

Direktlänk
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Annat format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annat språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf