liu.seSearch for publications in DiVA
Change search
Link to record
Permanent link

Direct link
Jonsson, Erik
Publications (10 of 17) Show all publications
Jonsson, E. & Felsberg, M. (2009). Efficient computation of channel-coded feature maps through piecewise polynomials. Image and Vision Computing, 27(11), 1688-1694
Open this publication in new window or tab >>Efficient computation of channel-coded feature maps through piecewise polynomials
2009 (English)In: Image and Vision Computing, ISSN 0262-8856, Vol. 27, no 11, p. 1688-1694Article in journal (Refereed) Published
Abstract [en]

Channel-coded feature maps (CCFMs) represent arbitrary image features using multi-dimensional histograms with soft and overlapping bins. This representation can be seen as a generalization of the SIFT descriptor, where one advantage is that it is better suited for computing derivatives with respect to image transformations. Using these derivatives, a local optimization of image scale, rotation and position relative to a reference view can be computed. If piecewise polynomial bin functions are used, e.g. B-splines, these histograms can be computed by first encoding the data set into a histogram-like representation with non-overlapping multi-dimensional monomials as bin functions. This representation can then be processed using multi-dimensional convolutions to obtain the desired representation. This allows to reuse much of the computations for the derivatives. By comparing the complexity of this method to direct encoding, it is found that the piecewise method is preferable for large images and smaller patches with few channels, which makes it useful, e.g. in early steps of coarse-to-fine approaches.

Keywords
Channel-coded feature maps; Feature histograms; Piecewise polynomials; Soft histograms; Splines
National Category
Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-21197 (URN)10.1016/j.imavis.2008.11.002 (DOI)
Note
Original Publication: Erik Jonsson and Michael Felsberg, Efficient computation of channel-coded feature maps through piecewise polynomials, 2009, Image and Vision Computing, (27), 11, 1688-1694. http://dx.doi.org/10.1016/j.imavis.2008.11.002 Copyright: Elsevier Science B.V., Amsterdam. http://www.elsevier.com/ Available from: 2009-09-30 Created: 2009-09-30 Last updated: 2016-05-04
Larsson, F., Jonsson, E. & Felsberg, M. (2009). Simultaneously learning to recognize and control a low-cost robotic arm. Image and Vision Computing, 27(11), 1729-1739
Open this publication in new window or tab >>Simultaneously learning to recognize and control a low-cost robotic arm
2009 (English)In: Image and Vision Computing, ISSN 0262-8856, E-ISSN 1872-8138, Vol. 27, no 11, p. 1729-1739Article in journal (Refereed) Published
Abstract [en]

In this paper, we present a visual servoing method based on a learned mapping between feature space and control space. Using a suitable recognition algorithm, we present and evaluate a complete method that simultaneously learns the appearance and control of a low-cost robotic arm. The recognition part is trained using an action precedes perception approach. The novelty of this paper, apart from the visual servoing method per se, is the combination of visual servoing with gripper recognition. We show that we can achieve high precision positioning without knowing in advance what the robotic arm looks like or how it is controlled.

Keywords
Gripper recognition; Jacobian estimation; LWPR; Visual servoing
National Category
Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-21195 (URN)10.1016/j.imavis.2009.04.003 (DOI)
Note
Original Publication: Fredrik Larsson, Erik Jonsson and Michael Felsberg, Simultaneously learning to recognize and control a low-cost robotic arm, 2009, Image and Vision Computing, (27), 11, 1729-1739. http://dx.doi.org/10.1016/j.imavis.2009.04.003 Copyright: Elsevier Science B.V., Amsterdam. http://www.elsevier.com/ Available from: 2009-09-30 Created: 2009-09-30 Last updated: 2017-12-13Bibliographically approved
Jonsson, E. (2008). Channel-Coded Feature Maps for Computer Vision and Machine Learning. (Doctoral dissertation). Institutionen för systemteknik
Open this publication in new window or tab >>Channel-Coded Feature Maps for Computer Vision and Machine Learning
2008 (English)Doctoral thesis, monograph (Other academic)
Abstract [en]

This thesis is about channel-coded feature maps applied in view-based object recognition, tracking, and machine learning. A channel-coded feature map is a soft histogram of joint spatial pixel positions and image feature values. Typical useful features include local orientation and color. Using these features, each channel measures the co-occurrence of a certain orientation and color at a certain position in an image or image patch. Channel-coded feature maps can be seen as a generalization of the SIFT descriptor with the options of including more features and replacing the linear interpolation between bins by a more general basis function.

The general idea of channel coding originates from a model of how information might be represented in the human brain. For example, different neurons tend to be sensitive to different orientations of local structures in the visual input. The sensitivity profiles tend to be smooth such that one neuron is maximally activated by a certain orientation, with a gradually decaying activity as the input is rotated.

This thesis extends previous work on using channel-coding ideas within computer vision and machine learning. By differentiating the channel-coded feature maps with respect to transformations of the underlying image, a method for image registration and tracking is constructed. By using piecewise polynomial basis functions, the channel coding can be computed more efficiently, and a general encoding method for N-dimensional feature spaces is presented.

Furthermore, I argue for using channel-coded feature maps in view-based pose estimation, where a continuous pose parameter is estimated from a query image given a number of training views with known pose. The optimization of position, rotation and scale of the object in the image plane is then included in the optimization problem, leading to a simultaneous tracking and pose estimation algorithm. Apart from objects and poses, the thesis examines the use of channel coding in connection with Bayesian networks. The goal here is to avoid the hard discretizations usually required when Markov random fields are used on intrinsically continuous signals like depth for stereo vision or color values in image restoration.

Channel coding has previously been used to design machine learning algorithms that are robust to outliers, ambiguities, and discontinuities in the training data. This is obtained by finding a linear mapping between channel-coded input and output values. This thesis extends this method with an incremental version and identifies and analyzes a key feature of the method -- that it is able to handle a learning situation where the correspondence structure between the input and output space is not completely known. In contrast to a traditional supervised learning setting, the training examples are groups of unordered input-output points, where the correspondence structure within each group is unknown. This behavior is studied theoretically and the effect of outliers and convergence properties are analyzed.

All presented methods have been evaluated experimentally. The work has been conducted within the cognitive systems research project COSPAL funded by EC FP6, and much of the contents has been put to use in the final COSPAL demonstrator system.

Place, publisher, year, edition, pages
Institutionen för systemteknik, 2008. p. 155
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 1160
Keywords
computer vision, machine learning, object recognition, pose estimation
National Category
Computer graphics and computer vision
Identifiers
urn:nbn:se:liu:diva-11040 (URN)978-91-7393-988-1 (ISBN)
Public defence
2008-03-28, Glashuset, Hus B, Campus Valla, Linköpings Universitet, Linköping, 13:15 (English)
Opponent
Supervisors
Available from: 2008-02-20 Created: 2008-02-20 Last updated: 2025-02-07
Larsson, F., Jonsson, E. & Felsberg, M. (2008). Learning Floppy Robot Control. In: SSBA,2008 (pp. 39-42).
Open this publication in new window or tab >>Learning Floppy Robot Control
2008 (English)In: SSBA,2008, 2008, p. 39-42Conference paper, Published paper (Other academic)
National Category
Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-44909 (URN)78193 (Local ID)78193 (Archive number)78193 (OAI)
Available from: 2009-10-10 Created: 2009-10-10 Last updated: 2016-05-04
Jonsson, E. (2008). Object Recognition using Channel-Coded Feature Maps: C++ Implementation Documentation (ed.). Linköping, Sweden: Linköping University, Department of Electrical Engineering
Open this publication in new window or tab >>Object Recognition using Channel-Coded Feature Maps: C++ Implementation Documentation
2008 (English)Report (Other academic)
Abstract [en]

This report gives an overview and motivates the design of a C++ framework for object recognition using channel-coded feature maps. The code was produced in connection to the work on my PhD thesis Channel-Coded Feature Maps for Object Recognition and Machine Learning. The package contains algorithms ranging from basic image processing routines to specific complex algorithms for creating channel-coded feature maps through piecewise polynomials. Much emphasis has been put in creating a flexible framework using virtual interfaces. This makes it easy e.g.~to switch between different image primitives detectors or learning methods in an object recognizer. Some common design choices include an image class with a convenient but fast pixel access, a configurable assert macro for error handling and a common base class for object ownership management. The main computer vision algorithms are channel-coded feature maps (CCFMs) including their derivatives, single-sided colored lines, object detection using an abstract hypothesize-verify framework and tracking and pose estimation using locally weighted regression and CCFMs. The code is considered as having alpha status at best. It is available under the GNU General Public License (GPL) and is mainly intended for future research on the subject.

Place, publisher, year, edition, pages
Linköping, Sweden: Linköping University, Department of Electrical Engineering, 2008. p. 17
Series
LiTH-ISY-R, ISSN 1400-3902 ; 2838
National Category
Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-53340 (URN)LITH-ISY-R-2838 (ISRN)
Available from: 2010-01-21 Created: 2010-01-20 Last updated: 2014-08-27Bibliographically approved
Jonsson, E. & Felsberg, M. (2007). Accurate Interpolation in Appearance-Based Pose Estimation. In: Svenska Sällskapet för Automatiserad Bildanalys SSBA Symposium,2007 (pp. 13-16).
Open this publication in new window or tab >>Accurate Interpolation in Appearance-Based Pose Estimation
2007 (English)In: Svenska Sällskapet för Automatiserad Bildanalys SSBA Symposium,2007, 2007, p. 13-16Conference paper, Published paper (Other academic)
National Category
Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-41620 (URN)58413 (Local ID)58413 (Archive number)58413 (OAI)
Available from: 2009-10-10 Created: 2009-10-10 Last updated: 2016-05-04
Jonsson, E. & Felsberg, M. (2007). Accurate Interpolation in Appearance-Based Pose Estimation. In: Bjarne Kjær Ersbøll and Kim Steenstrup Pedersen (Ed.), Bjarne Kjær Ersbøll and Kim Steenstrup Pedersen (Ed.), Image Analysis: 15th Scandinavian Conference, SCIA 2007, Aalborg, Denmark, June 10-14, 2007. Paper presented at 15th Scandinavian Conference, SCIA 2007, Aalborg, Denmark, June 10-14, 2007 (pp. 1-10). Springer Berlin/Heidelberg
Open this publication in new window or tab >>Accurate Interpolation in Appearance-Based Pose Estimation
2007 (English)In: Image Analysis: 15th Scandinavian Conference, SCIA 2007, Aalborg, Denmark, June 10-14, 2007 / [ed] Bjarne Kjær Ersbøll and Kim Steenstrup Pedersen, Springer Berlin/Heidelberg, 2007, p. 1-10Conference paper, Published paper (Refereed)
Abstract [en]

One problem in appearance-based pose estimation is the need for many training examples, i.e. images of the object in a large number of known poses. Some invariance can be obtained by considering translations, rotations and scale changes in the image plane, but the remaining degrees of freedom are often handled simply by sampling the pose space densely enough. This work presents a method for accurate interpolation between training views using local linear models. As a view representation local soft orientation histograms are used. The derivative of this representation with respect to the image plane transformations is computed, and a Gauss-Newton optimization is used to optimize all pose parameters simultaneously, resulting in an accurate estimate.

Place, publisher, year, edition, pages
Springer Berlin/Heidelberg, 2007
Series
Lecture Notes in Computer Science, ISSN 0302-9743, E-ISSN 1611-3349 ; 4522
National Category
Computer Sciences
Identifiers
urn:nbn:se:liu:diva-38253 (URN)10.1007/978-3-540-73040-8_1 (DOI)000247364000001 ()43285 (Local ID)978-3-540-73039-2 (ISBN)978-3-540-73040-8 (ISBN)43285 (Archive number)43285 (OAI)
Conference
15th Scandinavian Conference, SCIA 2007, Aalborg, Denmark, June 10-14, 2007
Available from: 2009-10-10 Created: 2009-10-10 Last updated: 2018-01-26Bibliographically approved
Felsberg, M., Wiklund, J., Jonsson, E., Moe, A. & Granlund, G. (2007). Exploratory Learning Strucutre in Artificial Cognitive Systems. In: International Cognitive Vision Workshop. Paper presented at The 5th International Conference on Computer Vision Systems, 2007, 21-24 March, Bielefeld University, Germany. Bielefeld: eCollections
Open this publication in new window or tab >>Exploratory Learning Strucutre in Artificial Cognitive Systems
Show others...
2007 (English)In: International Cognitive Vision Workshop, Bielefeld: eCollections , 2007Conference paper, Published paper (Other academic)
Abstract [en]

One major goal of the COSPAL project is to develop an artificial cognitive system architecture with the capability of exploratory learning. Exploratory learning is a strategy that allows to apply generalization on a conceptual level, resulting in an extension of competences. Whereas classical learning methods aim at best possible generalization, i.e., concluding from a number of samples of a problem class to the problem class itself, exploration aims at applying acquired competences to a new problem class. Incremental or online learning is an inherent requirement to perform exploratory learning.

Exploratory learning requires new theoretic tools and new algorithms. In the COSPAL project, we mainly investigate reinforcement-type learning methods for exploratory learning and in this paper we focus on its algorithmic aspect. Learning is performed in terms of four nested loops, where the outermost loop reflects the user-reinforcement-feedback loop, the intermediate two loops switch between different solution modes at symbolic respectively sub-symbolic level, and the innermost loop performs the acquired competences in terms of perception-action cycles. We present a system diagram which explains this process in more detail.

We discuss the learning strategy in terms of learning scenarios provided by the user. This interaction between user ('teacher') and system is a major difference to most existing systems where the system designer places his world model into the system. We believe that this is the key to extendable robust system behavior and successful interaction of humans and artificial cognitive systems.

We furthermore address the issue of bootstrapping the system, and, in particular, the visual recognition module. We give some more in-depth details about our recognition method and how feedback from higher levels is implemented. The described system is however work in progress and no final results are available yet. The available preliminary results that we have achieved so far, clearly point towards a successful proof of the architecture concept.

Place, publisher, year, edition, pages
Bielefeld: eCollections, 2007
Keywords
artificial cognitive system, perception action learning, exploratory learning, cognitive bootstrapping
National Category
Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-39511 (URN)10.2390/biecoll-icvs2007-173 (DOI)49069 (Local ID)49069 (Archive number)49069 (OAI)
Conference
The 5th International Conference on Computer Vision Systems, 2007, 21-24 March, Bielefeld University, Germany
Projects
COSPAL
Available from: 2009-10-10 Created: 2009-10-10 Last updated: 2016-05-04Bibliographically approved
Larsson, F., Jonsson, E. & Felsberg, M. (2007). Visual Servoing Based on Learned Inverse Kinematics. In: SSBA,2007 (pp. 21-24).
Open this publication in new window or tab >>Visual Servoing Based on Learned Inverse Kinematics
2007 (English)In: SSBA,2007, 2007, p. 21-24Conference paper, Published paper (Other academic)
National Category
Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-40820 (URN)54223 (Local ID)54223 (Archive number)54223 (OAI)
Available from: 2009-10-10 Created: 2009-10-10 Last updated: 2016-05-04
Larsson, F., Jonsson, E. & Felsberg, M. (2007). Visual Servoing for Floppy Robots using LWPR. In: RoboMat,2007.
Open this publication in new window or tab >>Visual Servoing for Floppy Robots using LWPR
2007 (English)In: RoboMat,2007, 2007Conference paper, Published paper (Refereed)
National Category
Engineering and Technology
Identifiers
urn:nbn:se:liu:diva-39508 (URN)49065 (Local ID)49065 (Archive number)49065 (OAI)
Available from: 2009-10-10 Created: 2009-10-10 Last updated: 2016-05-04
Organisations

Search in DiVA

Show all publications