liu.seSearch for publications in DiVA
Change search
Link to record
Permanent link

Direct link
Alternative names
Publications (10 of 196) Show all publications
Bae, C. & Gustafsson, O. (2025). Architectural Trade-Offs for High-Speed Real-Time Chromatic Dispersion Compensation FIR Filters With Time-Multiplexed Streaming FFTs. Journal of Lightwave Technology, 43(14), 6744-6753
Open this publication in new window or tab >>Architectural Trade-Offs for High-Speed Real-Time Chromatic Dispersion Compensation FIR Filters With Time-Multiplexed Streaming FFTs
2025 (English)In: Journal of Lightwave Technology, ISSN 0733-8724, E-ISSN 1558-2213, Vol. 43, no 14, p. 6744-6753Article in journal (Refereed) Published
Abstract [en]

Fiber-optic communication faces challenges to address impairment such as chromatic dispersion, which distorts signals, necessitating efficient chromatic dispersion compensation (CDC) solutions. In this work, architectural trade-offs when implementing finite-length impulse response (FIR) filters in the frequency-domain for high-speed CDC is presented. Most earlier works have considered a fully parallel FFT. However, as shown in this work, it is possible to reduce the power consumption by increasing the FFTs size, leading to time-multiplexed streaming FFT architectures. The results, implemented for a tentative 400G 16-QAM system, illustrates that traditionally used complexity measures, such as number of multiplications per sample, are not suitable when determining the best FFT size. It is also shown that this size is not necessarily a power of two. Instead, a ratio between the block length and the filter length of between one and two provides the best results in this case, lower than the commonly used multiplications per sample estimate. The results highlights the impact of the data shuffling, required for time-multiplexed FFTs, on the power consumption, but also the possibilities to lower the power consumption by not restricting to power-of-two FFTs.

Place, publisher, year, edition, pages
IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC, 2025
Keywords
Finite impulse response filters, Optical fiber filters, Frequency-domain analysis, Complexity theory, Clocks, Transforms, Optical fibers, Chromatic dispersion, Parallel processing, Optical signal processing
National Category
Signal Processing
Identifiers
urn:nbn:se:liu:diva-216638 (URN)10.1109/jlt.2025.3558510 (DOI)001530265700020 ()2-s2.0-105002488744 (Scopus ID)
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Available from: 2025-08-19 Created: 2025-08-19 Last updated: 2025-09-03
Lindberg, T. & Gustafsson, O. (2025). Exact Eight-Bit Floating-Point Multiplication Using Integer Arithmetic. In: 2025 59th Asilomar Conference on Signals, Systems, and Computers: . Paper presented at 2025 59th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, 26-29 Oct. 2025 (pp. 315-320). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Exact Eight-Bit Floating-Point Multiplication Using Integer Arithmetic
2025 (English)In: 2025 59th Asilomar Conference on Signals, Systems, and Computers, Institute of Electrical and Electronics Engineers (IEEE), 2025, p. 315-320Conference paper, Published paper (Refereed)
Abstract [en]

Exact eight-bit floating-point multiplication performed using simple integer operations is considered in this work. For two-bit mantissa formats, correct rounding for all considered rounding modes can be obtained, either directly or by adding a conditional carry in. Using slightly more complicated expressions for the carry in, all rounding modes except directed modes can be obtained for three-bit mantissa formats. Hardware implementation results using both standard-cell and FPGA technology are presented, illustrating the potential benefit of integer computation. Savings in terms of resources and delay are obtained on both platforms, especially for the three-bit mantissa formats. The approach can also be attractive for platforms with integer vector instructions.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2025
Series
Asilomar Conference on Signals, Systems & Computers, ISSN 1058-6393, E-ISSN 2576-2303
National Category
Other Medical Biotechnology
Identifiers
urn:nbn:se:liu:diva-223494 (URN)10.1109/ieeeconf67917.2025.11443802 (DOI)2-s2.0-105035836628 (Scopus ID)9798331587451 (ISBN)9798331587468 (ISBN)
Conference
2025 59th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, 26-29 Oct. 2025
Available from: 2026-05-04 Created: 2026-05-04 Last updated: 2026-05-04
Hansson, O., Nunez-Yanez, J. & Gustafsson, O. (2025). Finite Word-Length Effects for Symmetric Matrix Inversion with Vector-Based Algorithms. In: Proceedings of 59th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA: IEEE, 2025. Vol. 59, p. 294-299: . Paper presented at 2025 59th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, 26-29 October 2025 (pp. 294-299). Pacific Grove, CA, USA: IEEE, 59
Open this publication in new window or tab >>Finite Word-Length Effects for Symmetric Matrix Inversion with Vector-Based Algorithms
2025 (English)In: Proceedings of 59th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA: IEEE, 2025. Vol. 59, p. 294-299, Pacific Grove, CA, USA: IEEE, 2025, Vol. 59, p. 294-299Conference paper, Published paper (Refereed)
Abstract [en]

In this work, we present an investigation of the numerical properties of different symmetric matrix inversion algorithms, with a focus on two vector-based algorithms: Block LDLT and Tile LDLT. The inversion of symmetric positive-definite matrices finds applications in many different areas of signal processing, such as in MIMO detection and adaptive filtering. Using arithmetic simulations with the APyTypes Python software library, we demonstrate the numerical properties of different algorithms when using fixed-point number representations. The two algorithms in focus show potential, performing very close to the standard LDLT algorithm. The Tile LDLT algorithm slightly outperforms the Block LDLT algorithm in most cases, especially for larger matrix sizes. The application-specific hardware architecture for the two matrix inversion algorithms is only given at a high level, with the actual implementation left as future work.

Place, publisher, year, edition, pages
Pacific Grove, CA, USA: IEEE, 2025
Series
Asilomar Conference on Signals, Systems, and Computers, ISSN 1058-6393, E-ISSN 2576-2303 ; 59
Keywords
Matrix inversion, Symmetric matrix, Fixed-point, Vector-based algorithms, Block LDLT, Tile LDLT
National Category
Computer Engineering
Identifiers
urn:nbn:se:liu:diva-223466 (URN)10.1109/IEEECONF67917.2025.11443719 (DOI)2-s2.0-105035826152 (Scopus ID)9798331587451 (ISBN)9798331587468 (ISBN)
Conference
2025 59th Asilomar Conference on Signals, Systems, and Computers, Pacific Grove, CA, USA, 26-29 October 2025
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Available from: 2026-05-04 Created: 2026-05-04 Last updated: 2026-05-07
Svensson, L., Gustafsson, O. & Rodrigues, J. (2025). Scalable Low-latency Systolic Arrays using Toroidal and Bi-Directional Dataflows. In: IEEE Workshop on Signal Processing Systems (SIPS): . Paper presented at IEEE Workshop on Signal Processing Systems (SiPS), Hong Kong, 01-04 November 2025. IEEE
Open this publication in new window or tab >>Scalable Low-latency Systolic Arrays using Toroidal and Bi-Directional Dataflows
2025 (English)In: IEEE Workshop on Signal Processing Systems (SIPS), IEEE, 2025Conference paper, Published paper (Refereed)
Abstract [en]

Systolic arrays (SAs) for matrix multiplication are commonly used in machine learning (ML), wireless communication, and signal processing. Inherently offering high throughput with good data reuse, they are well-positioned for both low-power edge devices and accelerator applications in high-performance computing. Current realizations suffer from startup latency, defined as the time required to fully utilize all processing elements (PEs). In this work, this issue is addressed by introducing bidirectional systolic arrays with connected edges that form toroidal dataflows. The proposed systolic arrays significantly reduce computational and readout latency from 4n−2 to 2.5n−1 clock cycles for an n×n matrix multiplication, while simultaneously reducing energy per operation by up to 43% compared to conventional SAs. Moreover, a variety of differently shaped SAs are synthesized in a 22 nm CMOS technology, and it is shown that the toroidal designs offer a 5%−12% lower silicon area cost.

Place, publisher, year, edition, pages
IEEE, 2025
Series
2025 IEEE Workshop on Signal Processing Systems (SiPS), ISSN 1520-6130, E-ISSN 2374-7390
National Category
Computer Sciences Computer Systems
Identifiers
urn:nbn:se:liu:diva-223411 (URN)10.1109/SiPS66314.2025.11261266 (DOI)2-s2.0-105031673264 (Scopus ID)9798331598327 (ISBN)9798331598310 (ISBN)
Conference
IEEE Workshop on Signal Processing Systems (SiPS), Hong Kong, 01-04 November 2025
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications, D11
Available from: 2026-04-29 Created: 2026-04-29 Last updated: 2026-05-07
Henriksson, M., Lindberg, T. & Gustafsson, O. (2024). APyTypes: Algorithmic Data Types in Python for Efficient Simulation of Finite Word-Length Effects. In: Proceedings of the IEEE 31st Symposium on Computer Arithmetic (ARITH): . Paper presented at IEEE 31st Symposium on Computer Arithmetic (ARITH), Malaga, Spain (pp. 72-75). Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>APyTypes: Algorithmic Data Types in Python for Efficient Simulation of Finite Word-Length Effects
2024 (English)In: Proceedings of the IEEE 31st Symposium on Computer Arithmetic (ARITH), Institute of Electrical and Electronics Engineers (IEEE), 2024, p. 72-75Conference paper, Published paper (Refereed)
Abstract [en]

A new Python library, APyTypes, suitable for simulating and exploring finite word-length effects is presented. The library supports configurable bit-accurate fixed- and floating-point types of both scalars and multidimensional arrays and uses a C++ backend to accelerate runtime performance. The underlying design principles of the library are introduced and examples show how it can be used. We argue that APyTypes have significant advantages over existing arithmetic libraries, especially from a hardware design perspective. Finally, some directions for further work are outlined.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2024
National Category
Computer Systems
Identifiers
urn:nbn:se:liu:diva-205936 (URN)10.1109/ARITH61463.2024.00021 (DOI)001267256100021 ()2-s2.0-85198640146 (Scopus ID)9798350384321 (ISBN)9798350384338 (ISBN)
Conference
IEEE 31st Symposium on Computer Arithmetic (ARITH), Malaga, Spain
Note

Funding Agencies|ELLIIT strategic research environment through the project D3 ACRE - Approximate Computing Reducing Energy; Swedish Foundation for Strategic Research through the project Large Intelligent Surfaces - Architecture and Hardware

Available from: 2024-07-12 Created: 2024-07-12 Last updated: 2025-08-26
Lindberg, T., Henriksson, M. & Gustafsson, O. (2024). APyTypes: Flexible Fixed- and Floating-Point Types for Word Length Simulation in Python. In: 4th Workshop on Open-Source Design Automation (OSDA): . Paper presented at Open-Source Design Automation (OSDA).
Open this publication in new window or tab >>APyTypes: Flexible Fixed- and Floating-Point Types for Word Length Simulation in Python
2024 (English)In: 4th Workshop on Open-Source Design Automation (OSDA), 2024Conference paper, Published paper (Refereed)
Abstract [en]

APyTypes is a new Python library, written in C++, that provides scalars and multi-dimensional arrays with configurable fixed- and floating-point representations. It is suitable for finite word-length design exploration and to provide reference data for custom hardware implementations.

National Category
Computer Systems
Identifiers
urn:nbn:se:liu:diva-220676 (URN)
Conference
Open-Source Design Automation (OSDA)
Funder
ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications, D3Swedish Foundation for Strategic Research, CHI19-0001
Available from: 2026-04-30 Created: 2026-04-30 Last updated: 2026-05-07
Henriksson, M., Winbladh, H. & Gustafsson, O. (2024). Multi-Stream FFT Architectures for a Distributed MIMO Large Intelligent Surfaces Testbed. In: Nurmi J., Rodrigues J., Pezzarossa L., Aberg V., Behmanesh B. (Ed.), 2024 IEEE Nordic Circuits and Systems Conference, NORCAS 2024 - Proceedings: . Paper presented at IEEE Nordic Circuits and Systems Conference (NorCAS), Lund, Sweden, 29-30 October, 2024. Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Multi-Stream FFT Architectures for a Distributed MIMO Large Intelligent Surfaces Testbed
2024 (English)In: 2024 IEEE Nordic Circuits and Systems Conference, NORCAS 2024 - Proceedings / [ed] Nurmi J., Rodrigues J., Pezzarossa L., Aberg V., Behmanesh B., Institute of Electrical and Electronics Engineers (IEEE), 2024Conference paper, Published paper (Refereed)
Abstract [en]

A modern communication testbed based on distributed multiple-input multiple-output (MIMO) is constructed at Lund University. This testbed uses 16-antenna panels as a base for the distributed MIMO system. As the mixers in these panels are fully digital, there is a rational relationship between the sample frequency and the clock frequency of the field-programmable gate-array that performs the digital baseband processing. This rational relationship, combined with the multiple data streams in the panel, makes for an interesting design challenge when implementing high-utilization circuits. Three unique 16 multi-stream fast Fourier transform architectures tailored toward this particular distributed MIMO testbed are presented. The proposed architectures take the rational relationship between clock frequency and sample frequency into consideration and achieve 96.25%–100% butterfly utilization in the multi-streamFFT.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2024
Keywords
fast Fourier transform (FFT), field-programmable gate-array (FPGA), distributed multiple-input multiple-output (MIMO), multi-stream
National Category
Computer Engineering Signal Processing
Identifiers
urn:nbn:se:liu:diva-209881 (URN)10.1109/NorCAS64408.2024.10752445 (DOI)001444043400008 ()2-s2.0-85211892823 (Scopus ID)9798331517670 (ISBN)9798331517663 (ISBN)
Conference
IEEE Nordic Circuits and Systems Conference (NorCAS), Lund, Sweden, 29-30 October, 2024
Funder
Swedish Foundation for Strategic Research, CHI19-0001ELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Note

Funding Agencies|Swedish Foundation for Strategic Research (SSF) [CHI19-0001]; Excellence Center at Linkoping-Lund in Information Technology (ELLIIT)

Available from: 2024-11-19 Created: 2024-11-19 Last updated: 2025-11-14
Skarman, F. & Gustafsson, O. (2023). Abstraction in the Spade Hardware Description Language. In: : . Paper presented at LATTE ’23 - Workshop on Languages, Tools, and Techniques for Accelerator Design, Vancouver, BC, Canada, March 26, 2023.
Open this publication in new window or tab >>Abstraction in the Spade Hardware Description Language
2023 (English)Conference paper, Oral presentation only (Other academic)
Abstract [en]

Spade is an HDL that enhances the productivity of HDL designers byadding useful abstractions for hardware design. These abstractionsare zero- or low-cost, meaning that the designer still has full controlover what hardware gets generated.

Keywords
Hardware description languages, languages and compilers, Design automation
National Category
Computer Engineering
Identifiers
urn:nbn:se:liu:diva-198342 (URN)
Conference
LATTE ’23 - Workshop on Languages, Tools, and Techniques for Accelerator Design, Vancouver, BC, Canada, March 26, 2023
Available from: 2023-10-05 Created: 2023-10-05 Last updated: 2024-12-05Bibliographically approved
Khan, M. T. & Gustafsson, O. (2023). Analyzing Step-Size Approximation for Fixed-Point Implementation of LMS and BLMS Algorithms. In: 2023 IEEE Nordic Circuits and Systems Conference (NorCAS): . Paper presented at 31 October 2023 - 01 November 2023, Aalborg, Denmark. IEEE
Open this publication in new window or tab >>Analyzing Step-Size Approximation for Fixed-Point Implementation of LMS and BLMS Algorithms
2023 (English)In: 2023 IEEE Nordic Circuits and Systems Conference (NorCAS), IEEE, 2023Conference paper, Published paper (Refereed)
Abstract [en]

In this work, we analyze the step-size approximation for fixed-point least-mean-square (LMS) and block LMS (BLMS) algorithms. Our primary focus is on investigating how step size approximation impacts the convergence rate and steady-state mean square error (MSE) across varying block sizes and filter lengths. We consider three different FP quantized LMS and BLMS algorithms. The results demonstrate that the algorithm with two quantizers in single precision behaves approximately the same as one quantizer under quantized weights, regardless of block size and filter lengths. Subsequently, we explore the approximation effects of nearest power-of-two and their combinations with different design parameters on the convergence performance. Simulation results for within the context of a system identification problem under these approximations reveal intriguing insights. For instance, a single quantizer algorithm without quantized error is more robust than its counterpart under these approximations. Additionally, both single quantizer algorithms with combined power-of-two approximations matches the behavior of the actual step-size.

Place, publisher, year, edition, pages
IEEE, 2023
National Category
Signal Processing
Identifiers
urn:nbn:se:liu:diva-199293 (URN)10.1109/NorCAS58970.2023.10305481 (DOI)001103249500039 ()9798350337570 (ISBN)9798350337587 (ISBN)
Conference
31 October 2023 - 01 November 2023, Aalborg, Denmark
Available from: 2023-11-24 Created: 2023-11-24 Last updated: 2024-01-17Bibliographically approved
Skarman, F., Klemmer, L., Gustafsson, O. & Große, D. (2023). Enhancing Compiler-Driven HDL Design with Automatic Waveform Analysis. In: Forum on Specification, Verification and Design Languages, FDL: . Paper presented at Forum on Specification, Verification and Design Languages, FDL, 13-15 September 2023, Turin, Italy (pp. 1-8). IEEE
Open this publication in new window or tab >>Enhancing Compiler-Driven HDL Design with Automatic Waveform Analysis
2023 (English)In: Forum on Specification, Verification and Design Languages, FDL, IEEE, 2023, p. 1-8Conference paper, Published paper (Refereed)
Abstract [en]

The time-to-market of a new product is one of its most crucial factors for success, therefore, reducing this time is of utter importance. However, this reduction must not come at the expense of a less thorough development process. This paper presents a compiler-driven approach for automatically analyzing metrics such as transaction delays or bus throughput on simulation waveforms of projects developed in the Spade Hardware Description Language (HDL). By utilizing the Spade compiler’s knowledge about design internals, an automatic analysis of the waveforms created during simulation is possible using the Waveform Analysis Language (WAL). Analysis programs can be bundled with Spade projects or libraries, such that they are automatically detected by Spade and can be reused by other projects using simple annotations. We call these bundled WAL programs analysis passes, since they fit into the Spade workflow and provide thorough analysis at no additional cost to the users of these libraries. In a detailed description, we present how new analysis passes can be defined using the example of a data streaming interface. Additionally, we highlight the possibilities of analysis passes in two case studies, including Finite State Machine (FSM) and Wishbone protocol analysis.

Place, publisher, year, edition, pages
IEEE, 2023
Keywords
performance analysis, hardware description languages, debugging
National Category
Computer Sciences
Identifiers
urn:nbn:se:liu:diva-209367 (URN)10.1109/FDL59689.2023.10272204 (DOI)2-s2.0-85175238128 (Scopus ID)979-8-3503-0737-5 (ISBN)979-8-3503-0738-2 (ISBN)
Conference
Forum on Specification, Verification and Design Languages, FDL, 13-15 September 2023, Turin, Italy
Available from: 2024-11-11 Created: 2024-11-11 Last updated: 2025-11-17
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0003-3470-3911

Search in DiVA

Show all publications