liu.seSearch for publications in DiVA
Change search
Link to record
Permanent link

Direct link
Publications (5 of 5) Show all publications
Sikder, M. F., Ramachandranpillai, R., de Leng, D. & Heintz, F. (2025). Promoting Intersectional Fairness through Knowledge Distillation. In: Inês Lynce, Nello Murano, Mauro Vallati, Serena Villata, Federico Chesani, Michela Milano, Andrea Omicini, Mehdi Dastani (Ed.), : . Paper presented at 28th European Conference on Artificial Intelligence (ECAI), Bologna, Italy, 2025 (pp. 3427-3434). IOS Press
Open this publication in new window or tab >>Promoting Intersectional Fairness through Knowledge Distillation
2025 (English)In: / [ed] Inês Lynce, Nello Murano, Mauro Vallati, Serena Villata, Federico Chesani, Michela Milano, Andrea Omicini, Mehdi Dastani, IOS Press , 2025, p. 3427-3434Conference paper, Published paper (Refereed)
Abstract [en]

As Artificial Intelligence-driven decision-making systems become increasingly popular, ensuring fairness in their outcomes has emerged as a critical and urgent challenge. AI models, often trained on open-source datasets embedded with human and systemic biases, risk producing decisions that disadvantage certain demographics. This challenge intensifies when multiple sensitive attributes interact, leading to intersectional bias, a compounded and uniquely complex form of unfairness. Over the years, various methods have been proposed to address bias at the data and model levels. However, mitigating intersectional bias in decision-making remains an under-explored challenge. Motivated by this gap, we propose a novel framework that leverages knowledge distillation to promote intersectional fairness. Our approach proceeds in two stages: first, a teacher model is trained solely to maximize predictive accuracy, followed by a student model that inherits the teacher's representational knowledge while incorporating intersectional fairness constraints. The student model integrates tailored loss functions that enforce parity in false positive rates and demographic distributions across intersectional groups, alongside an adversarial objective that minimizes protected attribute information within the learned representation. Empirical evaluation across multiple benchmark datasets demonstrates that we achieve a 52% increase in accuracy for multi-class classification and a 61% reduction in average false positive rate across intersectional groups and outperforms state-of-the-art models. This distillation-based methodology provides a more stable optimization opportunity than direct fairness approaches, resulting in substantially fairer representations, particularly for multiple sensitive attributes and underrepresented demographic intersections.

Place, publisher, year, edition, pages
IOS Press, 2025
Keywords
Data Fairness, Representation Learning, Intersectional Fairness
National Category
Artificial Intelligence
Identifiers
urn:nbn:se:liu:diva-219032 (URN)10.3233/FAIA251214 (DOI)
Conference
28th European Conference on Artificial Intelligence (ECAI), Bologna, Italy, 2025
Funder
Knut and Alice Wallenberg FoundationELLIIT - The Linköping‐Lund Initiative on IT and Mobile Communications
Available from: 2025-10-25 Created: 2025-10-25 Last updated: 2025-10-30
Sikder, M. F. (2025). Representative Synthetic Data for Fair Decision Making. (Doctoral dissertation). Linköping: Linköping University Electronic Press
Open this publication in new window or tab >>Representative Synthetic Data for Fair Decision Making
2025 (English)Doctoral thesis, monograph (Other academic)
Abstract [en]

Deep generative models and representation learning techniques have become essential to modern machine learning, enabling both the creation of synthetic data and the extraction of meaningful features for downstream tasks. While synthetic data generation promises to address data scarcity concerns, and representation learning forms the backbone of decision-making systems, both approaches can amplify societal biases present in the training data. In high-stakes applications such as healthcare, criminal justice, and financial services, biased synthetic data can pass on discriminatory patterns to new datasets, while biased representation can lead to unfair decisions that disproportionately harm marginalized groups. The challenge becomes more extensive when dealing with intersectional bias, where discrimination occurs at the intersection of multiple sensitive attributes, e.g. race, gender, etc. In this dissertation, we propose approaches for fair synthetic data generation and fair representation learning that target bias mitigation for both individual sensitive attributes and their intersections.

To address model-specific biases in synthetic data generation, we first introduce Fair Latent Deep Generative Models (FLDGMs), a syntax-agnostic framework that first learns low-dimensional fair latent representations via a fairness-aware compression step, and then generates a synthetic fair latent space using either GANs or diffusion models, followed by high-fidelity reconstruction through a decoder. We also present Bias-transforming GAN (Bt-GAN), which tackles healthcare data biases by imposing information-theoretic constraints and preserves subgroup representation using density-aware sampling.

For fair representation, we develop two novel representation learning techniques specifically designed to address intersectional fairness. First, we present a knowledge distillation-based approach, where we distill knowledge from an accuracy-focused teacher into a student model that enforces intersectional fairness constraints, including False Positive Rate (FPR) and demographic parity, effectively reducing FPR disparities in multi-class settings. Second, Diff-Fair uses diffusion-based representation learning to minimize mutual information with sensitive attributes, integrating intersectional and FPR regularizers to reduce demographic and outcome disparities across subgroups, while maintaining strong accuracy in binary and multi-class tasks. Also, to enable systematic evaluation, we introduce FairX, an open-source benchmarking suite that integrates pre-, in-, and post-processing bias mitigation methods alongside fairness, utility, explainability, and synthetic-data evaluation metrics.

Along with fair synthetic data, we also investigate how to capture the distribution of time-series data using generative models. We present TransFusion, a diffusion and transformer-based architecture, that generates high-fidelity, long-sequenced synthetic time-series data (sequence length upto 384).

Together, these contributions advance the generation of synthetic data and fair representations by integrating bias-transforming models, intersectional fairness constraints, robust benchmarking, and long-sequence synthesis, offering practical tools for building more equitable AI systems.

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2025. p. 166
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 2495
National Category
Artificial Intelligence
Identifiers
urn:nbn:se:liu:diva-219906 (URN)10.3384/9789181183757 (DOI)9789181183740 (ISBN)9789181183757 (ISBN)
Public defence
2026-01-16, Ada Lovelace, B Building, Campus Valla, Linköping, 09:15 (English)
Opponent
Supervisors
Note

Funding agencies: This work was funded by the Knut and Alice Wallenberg Foundation, Sweden, the ELLIIT Excellence Center at Linköping-Lund for Information Technology, Sweden (portions of this work were carried out using the AIOps/Stellar), and TAILOR - an EU project with the aim to provide the scientific foundations for Trustworthy AI in Europe. The computations were enabled by the Berzelius resource provided by the Knut and Alice Wallenberg Foundation at the National Supercomputer Centre.

Available from: 2025-12-08 Created: 2025-12-08 Last updated: 2025-12-09Bibliographically approved
Ramachandranpillai, R., Sikder, M. F., Bergström, D. & Heintz, F. (2024). Bt-GAN: Generating Fair Synthetic Healthdata via Bias-transforming Generative Adversarial Networks. The journal of artificial intelligence research, 79, 1313-1341
Open this publication in new window or tab >>Bt-GAN: Generating Fair Synthetic Healthdata via Bias-transforming Generative Adversarial Networks
2024 (English)In: The journal of artificial intelligence research, ISSN 1076-9757, E-ISSN 1943-5037, Vol. 79, p. 1313-1341Article in journal (Refereed) Published
Abstract [en]

Synthetic data generation offers a promising solution to enhance the usefulness of Electronic Healthcare Records (EHR) by generating realistic de-identified data. However, the existing literature primarily focuses on the quality of synthetic health data, neglecting the crucial aspect of fairness in downstream predictions. Consequently, models trained on synthetic EHR have faced criticism for producing biased outcomes in target tasks. These biases can arise from either spurious correlations between features or the failure of models to accurately represent sub-groups. To address these concerns, we present Bias-transforming Generative Adversarial Networks (Bt-GAN), a GAN-based synthetic data generator specifically designed for the healthcare domain. In order to tackle spurious correlations (i), we propose an information-constrained Data Generation Process (DGP) that enables the generator to learn a fair deterministic transformation based on a well-defined notion of algorithmic fairness. To overcome the challenge of capturing exact sub-group representations (ii), we incentivize the generator to preserve sub-group densities through score-based weighted sampling. This approach compels the generator to learn from underrepresented regions of the data manifold. To evaluate the effectiveness of our proposed method, we conduct extensive experiments using the Medical Information Mart for Intensive Care (MIMIC-III) database. Our results demonstrate that Bt-GAN achieves state-of-the-art accuracy while significantly improving fairness and minimizing bias amplification. Furthermore, we perform an in-depth explainability analysis to provide additional evidence supporting the validity of our study. In conclusion, our research introduces a novel and professional approach to addressing the limitations of synthetic data generation in the healthcare domain. By incorporating fairness considerations and leveraging advanced techniques such as GANs, we pave the way for more reliable and unbiased predictions in healthcare applications.

Place, publisher, year, edition, pages
AAAI Press, 2024
Keywords
Fair data generation, Trustworthy AI, Synthetic data generation, MIMIC-III, EHR
National Category
Computer Systems
Identifiers
urn:nbn:se:liu:diva-203151 (URN)10.1613/jair.1.15317 (DOI)001218386100001 ()
Note

Funding Agencies|Knut and Alice Wallenberg Foundation; ELLIIT Excellence Center at Linkoeping-Lund for Information Technology; TAILOR-an EU project

Available from: 2024-04-30 Created: 2024-04-30 Last updated: 2025-03-30Bibliographically approved
Sikder, M. F., Ramachandranpillai, R., de Leng, D. & Heintz, F. (2024). FairX: A comprehensive benchmarking tool for model analysis using fairness, utility, and explainability. In: Roberta Calegari,Virginia Dignum, Barry O'Sullivan (Ed.), Proceedings of the 2nd Workshop on Fairness and Bias in AI, co-located with 27th European Conference on Artificial Intelligence (ECAI 2024): . Paper presented at 2nd Workshop on Fairness and Bias in AI (AEQUITAS), co-located with 27th European Conference on Artificial Intelligence (ECAI 2024). CEUR, 3808, Article ID 16.
Open this publication in new window or tab >>FairX: A comprehensive benchmarking tool for model analysis using fairness, utility, and explainability
2024 (English)In: Proceedings of the 2nd Workshop on Fairness and Bias in AI, co-located with 27th European Conference on Artificial Intelligence (ECAI 2024) / [ed] Roberta Calegari,Virginia Dignum, Barry O'Sullivan, CEUR , 2024, Vol. 3808, article id 16Conference paper, Published paper (Refereed)
Abstract [en]

We present FairX, an open-source Python-based benchmarking tool designed for the comprehensive analysis of models under the umbrella of fairness, utility, and eXplainability (XAI). FairX enables users to train benchmarking bias-mitigation models and evaluate their fairness using a wide array of fairness metrics, data utility metrics, and generate explanations for model predictions, all within a unified framework. Existing benchmarking tools do not have the way to evaluate synthetic data generated from fair generative models, also they do not have the support for training fair generative models either. In FairX, we add fair generative models in the collection of our fair-model library (pre-processing, in-processing, post-processing) and evaluation metrics for evaluating the quality of synthetic fair data. This version of FairX supports both tabular and image datasets. It also allows users to provide their own custom datasets. The open-source FairX benchmarking package is publicly available at https://github.com/fahim-sikder/FairX.

Place, publisher, year, edition, pages
CEUR, 2024
Series
CEUR Workshop Proceedings, ISSN 1613-0073
Keywords
Data Fairness, Benchmarking, Synthetic Data, Evaluation
National Category
Computer Sciences
Identifiers
urn:nbn:se:liu:diva-209224 (URN)2-s2.0-85209988687 (Scopus ID)
Conference
2nd Workshop on Fairness and Bias in AI (AEQUITAS), co-located with 27th European Conference on Artificial Intelligence (ECAI 2024)
Funder
Knut and Alice Wallenberg Foundation
Available from: 2024-11-06 Created: 2024-11-06 Last updated: 2025-11-03Bibliographically approved
Sikder, M. F., Ferdous, M., Afroz, S., Podder, U., Fatema, K., Hossain, M. N., . . . Baowaly, M. K. (2023). Explainable Bengali Multiclass News Classification. In: 2023 26th International Conference on Computer and Information Technology (ICCIT): . Paper presented at 26th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 13-15 December 2023.. Institute of Electrical and Electronics Engineers (IEEE)
Open this publication in new window or tab >>Explainable Bengali Multiclass News Classification
Show others...
2023 (English)In: 2023 26th International Conference on Computer and Information Technology (ICCIT), Institute of Electrical and Electronics Engineers (IEEE), 2023Conference paper, Published paper (Refereed)
Abstract [en]

The automatic classification of news articles is crucial in the era of information overflow as it assists readers in accessing relevant information in a timely manner. Even though text classification is not a new area of study, there is potential for advancement concerning the Bengali language. Unlike other languages, Bengali is a complex language, and most of the datasets available online are imbalanced in terms of class label distribution. To increase the performance of classification methods and make them robust to handle imbalanced data, in this work, we propose a model consisting of pre-trained BERT architecture. We use a publicly available dataset of Bengali news articles with nine classes and achieve 92% accuracy. Along with the classification, explaining the model and the result is necessary for the application of trustworthy Artificial Intelligence. From this motivation, we use Integrated Gradient, an explainable AI technique, to explain the outcome of our model. We show which words in a news article affect the model to choose a particular class.

Place, publisher, year, edition, pages
Institute of Electrical and Electronics Engineers (IEEE), 2023
Keywords
BERT, Text Classification, Bengali News, Trustworthy AI, Explainability
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-201246 (URN)10.1109/ICCIT60459.2023.10441218 (DOI)9798350359015 (ISBN)9798350359022 (ISBN)
Conference
26th International Conference on Computer and Information Technology (ICCIT), Cox's Bazar, Bangladesh, 13-15 December 2023.
Available from: 2024-02-28 Created: 2024-02-28 Last updated: 2025-02-07Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0001-5307-997X

Search in DiVA

Show all publications

Profile pages

Personal site