liu.seSearch for publications in DiVA
Change search
Link to record
Permanent link

Direct link
Arpteg, Anders
Publications (2 of 2) Show all publications
Arpteg, A. (2005). Intelligent semi-structured information extraction: a user-driven approach to information extraction. (Doctoral dissertation). Linköping: Linköping University Electronic Press
Open this publication in new window or tab >>Intelligent semi-structured information extraction: a user-driven approach to information extraction
2005 (English)Doctoral thesis, monograph (Other academic)
Abstract [en]

The number of domains and tasks where information extraction tools can be used needs to be increased. One way to reach this goal is to design user-driven information extraction systems where non-expert users are able to adapt them to new domains and tasks. It is difficult to design general extraction systems that do not require expert skills or a large amount of work from the user. Therefore, it is difficult to increase the number of domains and tasks. A possible alternative is to design user-driven systems, which solve that problem by letting a large number of non-expert users adapt the systems themselves. To accomplish this goal, the systems need to become more intelligent and able to learn to extract with as little given information as possible.

The type of information extraction system that is in focus for this thesis is semi-structured information extraction. The term semi-structured refers to documents that not only contain natural language text but also additional structural information. The typical application is information extraction from World Wide Web hypertext documents. By making effective use of not only the link structure but also the structural information within each such document, user-driven extraction systems with high performance can be built.

There are two different approaches presented in this thesis to solve the user-driven extraction problem. The first takes a machine learning approach and tries to solve the problem using a modified $Q(\lambda)$ reinforcement learning algorithm. A problem with the first approach was that it was difficult to handle extraction from the hidden Web. Since the hidden Web is about 500 times larger than the visible Web, it would be very useful to be able to extract information from that part of the Web as well. The second approach is called the hidden observation approach and tries to also solve the problem of extracting from the hidden Web. The goal is to have a user-driven information extraction system that is also able to handle the hidden Web. The second approach uses a large part of the system developed for the first approach, but the additional information that is silently obtained from the user presents other problems and possibilities.

An agent-oriented system was designed to evaluate the approaches presented in this thesis. A set of experiments was conducted and the results indicate that a user-driven information extraction system is possible and no longer just a concept. However, additional work and research is necessary before a fully-fledged user-driven system can be designed.

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2005. p. 139
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 946
Keywords
Artificiell intelligens, Informaitonsåtervinning, Artificial intelligence
National Category
Computer Sciences
Identifiers
urn:nbn:se:liu:diva-33098 (URN)19075 (Local ID)91-85297-98-4 (ISBN)19075 (Archive number)19075 (OAI)
Public defence
2005-05-20, Key 1, Hus Key, Campus Valla, Linköpings universitet, Linköping, 10:15 (English)
Note

This work has been supported by University of Kalmar and the Knowledge Foundation.

Available from: 2009-10-09 Created: 2009-10-09 Last updated: 2018-01-13
Arpteg, A. (2003). Adaptive Semi-structured Information Extraction. (Licentiate dissertation). Institutionen för datavetenskap
Open this publication in new window or tab >>Adaptive Semi-structured Information Extraction
2003 (English)Licentiate thesis, monograph (Other academic)
Abstract [en]

The number of domains and tasks where information extraction tools can be used needs to be increased. One way to reach this goal is to construct user-driven information extraction systems where novice users are able to adapt them to new domains and tasks. To accomplish this goal, the systems need to become more intelligent and able to learn to extract information without need of expert skills or time-consuming work from the user.

The type of information extraction system that is in focus for this thesis is semistructural information extraction. The term semi-structural refers to documents that not only contain natural language text but also additional structural information. The typical application is information extraction from World Wide Web hypertext documents. By making effective use of not only the link structure but also the structural information within each such document, user-driven extraction systems with high performance can be built.

The extraction process contains several steps where different types of techniques are used. Examples of such types of techniques are those that take advantage of structural, pure syntactic, linguistic, and semantic information. The first step that is in focus for this thesis is the navigation step that takes advantage of the structural information. It is only one part of a complete extraction system, but it is an important part. The use of reinforcement learning algorithms for the navigation step can make the adaptation of the system to new tasks and domains more user-driven. The advantage of using reinforcement learning techniques is that the extraction agent can efficiently learn from its own experience without need for intensive user interactions.

An agent-oriented system was designed to evaluate the approach suggested in this thesis. Initial experiments showed that the training of the navigation step and the approach of the system was promising. However, additional components need to be included in the system before it becomes a fully-fledged user-driven system.

Place, publisher, year, edition, pages
Institutionen för datavetenskap, 2003. p. 85
Series
Linköping Studies in Science and Technology. Thesis, ISSN 0280-7971 ; 1000
Keywords
Information extraction, Artificial intelligence, Semi-structured data, Reinforced learning, Knowledge management
National Category
Computer Sciences
Identifiers
urn:nbn:se:liu:diva-5688 (URN)LiU-Tek-Lic-2002:73 (Local ID)9173735892 (ISBN)LiU-Tek-Lic-2002:73 (Archive number)LiU-Tek-Lic-2002:73 (OAI)
Presentation
2002-12-15, 00:00 (English)
Supervisors
Note

Report code: LiU-Tek-Lic-2002:73.

Available from: 2003-01-30 Created: 2003-01-30 Last updated: 2023-01-25Bibliographically approved
Organisations

Search in DiVA

Show all publications