liu.seSearch for publications in DiVA
Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Adaptive Semi-structured Information Extraction
Linköpings universitet, Institutionen för datavetenskap, KPLAB - Laboratoriet för kunskapsbearbetning. Linköpings universitet, Tekniska högskolan.
2003 (engelsk)Licentiatavhandling, monografi (Annet vitenskapelig)
Abstract [en]

The number of domains and tasks where information extraction tools can be used needs to be increased. One way to reach this goal is to construct user-driven information extraction systems where novice users are able to adapt them to new domains and tasks. To accomplish this goal, the systems need to become more intelligent and able to learn to extract information without need of expert skills or time-consuming work from the user.

The type of information extraction system that is in focus for this thesis is semistructural information extraction. The term semi-structural refers to documents that not only contain natural language text but also additional structural information. The typical application is information extraction from World Wide Web hypertext documents. By making effective use of not only the link structure but also the structural information within each such document, user-driven extraction systems with high performance can be built.

The extraction process contains several steps where different types of techniques are used. Examples of such types of techniques are those that take advantage of structural, pure syntactic, linguistic, and semantic information. The first step that is in focus for this thesis is the navigation step that takes advantage of the structural information. It is only one part of a complete extraction system, but it is an important part. The use of reinforcement learning algorithms for the navigation step can make the adaptation of the system to new tasks and domains more user-driven. The advantage of using reinforcement learning techniques is that the extraction agent can efficiently learn from its own experience without need for intensive user interactions.

An agent-oriented system was designed to evaluate the approach suggested in this thesis. Initial experiments showed that the training of the navigation step and the approach of the system was promising. However, additional components need to be included in the system before it becomes a fully-fledged user-driven system.

sted, utgiver, år, opplag, sider
Institutionen för datavetenskap , 2003. , s. 85
Serie
Linköping Studies in Science and Technology. Thesis, ISSN 0280-7971 ; 1000
Emneord [en]
Information extraction, Artificial intelligence, Semi-structured data, Reinforced learning, Knowledge management
HSV kategori
Identifikatorer
URN: urn:nbn:se:liu:diva-5688Lokal ID: LiU-Tek-Lic-2002:73ISBN: 9173735892 (tryckt)OAI: oai:DiVA.org:liu-5688DiVA, id: diva2:21449
Presentation
2002-12-15, 00:00 (engelsk)
Veileder
Merknad

Report code: LiU-Tek-Lic-2002:73.

Tilgjengelig fra: 2003-01-30 Laget: 2003-01-30 Sist oppdatert: 2023-01-25bibliografisk kontrollert

Open Access i DiVA

fulltekst(432 kB)1437 nedlastinger
Filinformasjon
Fil FULLTEXT01.pdfFilstørrelse 432 kBChecksum SHA-1
e5702c06b15b52e7b4564f08fefaa8cf1e2c3fd4390956952340eec4d949469ef3da7d1f
Type fulltextMimetype application/pdf

Person

Arpteg, Anders

Søk i DiVA

Av forfatter/redaktør
Arpteg, Anders
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar
Totalt: 1443 nedlastinger
Antall nedlastinger er summen av alle nedlastinger av alle fulltekster. Det kan for eksempel være tidligere versjoner som er ikke lenger tilgjengelige

isbn
urn-nbn

Altmetric

isbn
urn-nbn
Totalt: 1400 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • oxford
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf