liu.seSearch for publications in DiVA
Change search
Link to record
Permanent link

Direct link
Kunz, Jenny
Publications (10 of 11) Show all publications
Oji, R. & Kunz, J. (2025). How to Tune a Multilingual Encoder Model for Germanic Languages: A Study of PEFT, Full Fine-Tuning, and Language Adapters. In: Richard Johansson, Sara Stymne (Ed.), Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025): . Paper presented at The Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025), Tallinn, Estonia, March 3 – 4, 2025 (pp. 433-439). Tallinn, Estonia: University of Tartu Library
Open this publication in new window or tab >>How to Tune a Multilingual Encoder Model for Germanic Languages: A Study of PEFT, Full Fine-Tuning, and Language Adapters
2025 (English)In: Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025) / [ed] Richard Johansson, Sara Stymne, Tallinn, Estonia: University of Tartu Library , 2025, p. 433-439Conference paper, Published paper (Refereed)
Abstract [en]

This paper investigates the optimal use of the multilingual encoder model mDeBERTa for tasks in three Germanic languages – German, Swedish, and Icelandic – representing varying levels of presence and likely data quality in mDeBERTas pre-training data. We compare full finetuning with the parameter-efficient finetuning (PEFT) methods LoRA and Pfeiffer bottleneck adapters, finding that PEFT is more effective for the higher-resource language, German. However, results for Swedish and Icelandic are less consistent. We also observe differences between tasks: While PEFT tends to work better for question answering, full fine-tuning is preferable for named entity recognition. Inspired by previous research on modular approaches that combine task and language adapters, we evaluate the impact of adding PEFT modules trained on unstructured text, finding that this approach is not beneficial.

Place, publisher, year, edition, pages
Tallinn, Estonia: University of Tartu Library, 2025
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-219881 (URN)
Conference
The Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025), Tallinn, Estonia, March 3 – 4, 2025
Funder
CUGS (National Graduate School in Computer Science)Wallenberg AI, Autonomous Systems and Software Program (WASP)
Available from: 2025-12-04 Created: 2025-12-04 Last updated: 2026-06-24
Schlenker, J., Kunz, J., Anikina, T., Neumann, G. & Ostermann, S. (2025). Only for the Unseen Languages, Say the Llamas: On the Efficacy of Language Adapters for Cross-lingual Transfer in English-centric LLMs. In: Zhao, Jin., Wang, Mingyang., Liu, Zhu" (Ed.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop): . Paper presented at 63rd Association for Computational Linguistics Meeting-ACL-Annual, Vienna, Austria, Jul 27-Aug 01, 2025 (pp. 849-871). Association for Computational Linguistics
Open this publication in new window or tab >>Only for the Unseen Languages, Say the Llamas: On the Efficacy of Language Adapters for Cross-lingual Transfer in English-centric LLMs
Show others...
2025 (English)In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop) / [ed] Zhao, Jin., Wang, Mingyang., Liu, Zhu", Association for Computational Linguistics, 2025, p. 849-871Conference paper, Published paper (Refereed)
Abstract [en]

Most state-of-the-art large language models (LLMs) are trained mainly on English data, limiting their effectiveness on non-English, especially low-resource, languages. This study investigates whether language adapters can facilitate cross-lingual transfer in English-centric LLMs. We train language adapters for 13 languages using Llama 2 (7B) and Llama 3.1 (8B) as base models, and evaluate their effectiveness on two downstream tasks (MLQA and SIB-200) using either task adapters or in-context learning. Our results reveal that language adapters improve performance for languages not seen during pretraining, but provide negligible benefit for seen languages. These findings highlight the limitations of language adapters as a general solution for multilingual adaptation in English-centric LLMs.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2025
Keywords
large language models, language adaptation, parameter-efficient fine-tuning
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-219975 (URN)10.18653/v1/2025.acl-srw.62 (DOI)001596032700058 ()2-s2.0-105020387062 (Scopus ID)9798891762541 (ISBN)
Conference
63rd Association for Computational Linguistics Meeting-ACL-Annual, Vienna, Austria, Jul 27-Aug 01, 2025
Note

Funding Agencies|German Federal Ministry of Research, Technology and Space (BMFTR) [01IW24005]; Horizon Europe [101079164]

Available from: 2025-12-13 Created: 2025-12-13 Last updated: 2026-06-17
Kunz, J. (2025). Train More Parameters But Mind Their Placement: Insights into Language Adaptation with PEFT. In: Richard Johansson, Sara Stymne (Ed.), Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025): . Paper presented at Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025) (pp. 323-330). Tallinn, Estonia: University of Tartu Library, 25, Article ID 35.
Open this publication in new window or tab >>Train More Parameters But Mind Their Placement: Insights into Language Adaptation with PEFT
2025 (English)In: Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025) / [ed] Richard Johansson, Sara Stymne, Tallinn, Estonia: University of Tartu Library , 2025, Vol. 25, p. 323-330, article id 35Conference paper, Published paper (Refereed)
Abstract [en]

 Smaller LLMs still face significant challenges even in medium-resourced languages, particularly when it comes to language-specific knowledge – a problem not easily resolved with machine-translated data. In this case study on Icelandic, we aim to enhance the generation performance of an LLM by specialising it using unstructured text corpora. A key focus is on preventing interference with the models’ capabilities of handling longer context during this adaptation. Through ablation studies using various parameter-efficient fine-tuning (PEFT) methods and setups, we find that increasing the number of trainable parameters leads to better and more robust language adaptation. LoRAs placed in the feed-forward layers and bottleneck adapters show promising results with sufficient parameters, while prefix tuning and (IA)3 are not suitable. Although improvements are consistent in 0-shot summarisation, some adapted models struggle with longer context lengths, an issue that can be mitigated by adapting only the final layers.

Place, publisher, year, edition, pages
Tallinn, Estonia: University of Tartu Library, 2025
Keywords
Large language models, language adaptation, parameter-efficient fine-tuning
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-219974 (URN)
Conference
Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025)
Available from: 2025-12-13 Created: 2025-12-13 Last updated: 2026-06-24
Braun, M. & Kunz, J. (2024). A Hypothesis-Driven Framework for the Analysis of Self-Rationalising Models. In: Falk N., Papi S., Zhang M. (Ed.), Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop: . Paper presented at P18th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2024 - Student Research Workshop, SRW 2024 St. Julian's 21 March 2024 through 22 March 2024 (pp. 148-161).
Open this publication in new window or tab >>A Hypothesis-Driven Framework for the Analysis of Self-Rationalising Models
2024 (English)In: Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop / [ed] Falk N., Papi S., Zhang M., 2024, p. 148-161Conference paper, Published paper (Refereed)
Abstract [en]

The self-rationalising capabilities of LLMs are appealing because the generated explanations can give insights into the plausibility of the predictions. However, how faithful the explanations are to the predictions is questionable, raising the need to explore the patterns behind them further. To this end, we propose a hypothesis-driven statistical framework. We use a Bayesian network to implement a hypothesis about how a task (in our example, natural language inference) is solved, and its internal states are translated into natural language with templates. Those explanations are then compared to LLM-generated free-text explanations using automatic and human evaluations. This allows us to judge how similar the LLM’s and the Bayesian network’s decision processes are. We demonstrate the usage of our framework with an example hypothesis and two realisations in Bayesian networks. The resulting models do not exhibit a strong similarity to GPT-3.5. We discuss the implications of this as well as the framework’s potential to approximate LLM decisions better in future work.

Keywords
Computational linguistics
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-201911 (URN)001356735000011 ()2-s2.0-85188732502 (Scopus ID)9798891760905 (ISBN)
Conference
P18th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2024 - Student Research Workshop, SRW 2024 St. Julian's 21 March 2024 through 22 March 2024
Available from: 2024-03-26 Created: 2024-03-26 Last updated: 2025-02-07
Kunz, J. & Kuhlmann, M. (2024). Properties and Challenges of LLM-Generated Explanations. In: Su Lin Blodgett, Amanda Cercas Curry, Sunipa Dey, Michael Madaio, Ani Nenkova, Diyi Yang, Ziang Xiao (Ed.), Proceedings of the Third Workshop on Bridging Human-Computer Interaction and Natural Language Processing: . Paper presented at Third Workshop on Bridging Human-Computer Interaction and Natural Language Processing at NAACL 2024, MEXICO, JUN 21, 2024 (pp. 13-27). ASSOC COMPUTATIONAL LINGUISTICS-ACL
Open this publication in new window or tab >>Properties and Challenges of LLM-Generated Explanations
2024 (English)In: Proceedings of the Third Workshop on Bridging Human-Computer Interaction and Natural Language Processing / [ed] Su Lin Blodgett, Amanda Cercas Curry, Sunipa Dey, Michael Madaio, Ani Nenkova, Diyi Yang, Ziang Xiao, ASSOC COMPUTATIONAL LINGUISTICS-ACL , 2024, p. 13-27Conference paper, Published paper (Refereed)
Abstract [en]

The self-rationalising capabilities of large language models (LLMs) have been explored in restricted settings, using task-specific data sets.However, current LLMs do not (only) rely on specifically annotated data; nonetheless, they frequently explain their outputs.The properties of the generated explanations are influenced by the pre-training corpus and by the target data used for instruction fine-tuning.As the pre-training corpus includes a large amount of human-written explanations “in the wild”, we hypothesise that LLMs adopt common properties of human explanations.By analysing the outputs for a multi-domain instruction fine-tuning data set, we find that generated explanations show selectivity and contain illustrative elements, but less frequently are subjective or misleading.We discuss reasons and consequences of the properties’ presence or absence. In particular, we outline positive and negative implications depending on the goals and user groups of the self-rationalising system.

Place, publisher, year, edition, pages
ASSOC COMPUTATIONAL LINGUISTICS-ACL, 2024
Keywords
Natural Language Processing, Large Language Models, Explainability, Human-AI Interaction
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-207112 (URN)10.18653/v1/2024.hcinlp-1.2 (DOI)001596065400002 ()9798891761117 (ISBN)
Conference
Third Workshop on Bridging Human-Computer Interaction and Natural Language Processing at NAACL 2024, MEXICO, JUN 21, 2024
Note

Funding Agencies|Wallenberg AI, Autonomous Systems and Software Program (WASP) - Knut and Alice Wallenberg Foundation; European Commission [101135671]

Available from: 2024-09-02 Created: 2024-09-02 Last updated: 2025-12-18
Kunz, J. & Holmström, O. (2024). The Impact of Language Adapters in Cross-Lingual Transfer for NLU. In: Raúl Vázquez, Timothee Mickus, Jörg Tiedemann, Ivan Vulić, Ahmet Üstün (Ed.), Proceedings of the 1st Workshop on Modular and Open Multilingual NLP (MOOMIN 2024): . Paper presented at 1st Workshop on Modular and Open Multilingual NLP, MOOMIN 2024 St. Julian's 21 March 2024 (pp. 24-43). Association for Computational Linguistics
Open this publication in new window or tab >>The Impact of Language Adapters in Cross-Lingual Transfer for NLU
2024 (English)In: Proceedings of the 1st Workshop on Modular and Open Multilingual NLP (MOOMIN 2024) / [ed] Raúl Vázquez, Timothee Mickus, Jörg Tiedemann, Ivan Vulić, Ahmet Üstün, Association for Computational Linguistics , 2024, p. 24-43Conference paper, Published paper (Refereed)
Abstract [en]

Modular deep learning has been proposed for the efficient adaption of pre-trained models to new tasks, domains and languages. In particular, combining language adapters with task adapters has shown potential where no supervised data exists for a language. In this paper, we explore the role of language adapters in zero-shot cross-lingual transfer for natural language understanding (NLU) benchmarks. We study the effect of including a target-language adapter in detailed ablation studies with two multilingual models and three multilingual datasets. Our results show that the effect of target-language adapters is highly inconsistent across tasks, languages and models. Retaining the source-language adapter instead often leads to an equivalent, and sometimes to a better, performance. Removing the language adapter after training has only a weak negative effect, indicating that the language adapters do not have a strong impact on the predictions.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2024
Keywords
Large Language Models, LLMs, Adapters, NLP
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-201912 (URN)2-s2.0-85188800351 (Scopus ID)979-8-89176-084-4 (ISBN)
Conference
1st Workshop on Modular and Open Multilingual NLP, MOOMIN 2024 St. Julian's 21 March 2024
Available from: 2024-03-26 Created: 2024-03-26 Last updated: 2025-02-07
Kunz, J. (2024). Understanding Large Language Models: Towards Rigorous and Targeted Interpretability Using Probing Classifiers and Self-Rationalisation. (Doctoral dissertation). Linköping: Linköping University Electronic Press
Open this publication in new window or tab >>Understanding Large Language Models: Towards Rigorous and Targeted Interpretability Using Probing Classifiers and Self-Rationalisation
2024 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Large language models (LLMs) have become the base of many natural language processing (NLP) systems due to their performance and easy adaptability to various tasks. However, much about their inner workings is still unknown. LLMs have many millions or billions of parameters, and large parts of their training happen in a self-supervised fashion: They simply learn to predict the next word, or missing words, in a sequence. This is effective for picking up a wide range of linguistic, factual and relational information, but it implies that it is not trivial what exactly is learned, and how it is represented within the LLM. 

In this thesis, I present our work on methods contributing to better understanding LLMs. The work can be grouped into two approaches. The first lies within the field of interpretability, which is concerned with understanding the internal workings of the LLMs. Specifically, we analyse and refine a tool called probing classifiers that inspects the intermediate representations of LLMs, focusing on what roles the various layers of the neural model play. This helps us to get a global understanding of how information is structured in the model. I present our work on assessing and improving the probing methodologies. We developed a framework to clarify the limitations of past methods, showing that all common controls are insufficient. Based on this, we proposed more restrictive probing setups by creating artificial distribution shifts. We developed new metrics for the evaluation of probing classifiers that move the focus from the overall information that the layer contains to differences in information content across the LLM. 

The second approach is concerned with explainability, specifically with self-rationalising models that generate free-text explanations along with their predictions. This is an instance of local understandability: We obtain justifications for individual predictions. In this setup, however, the generation of the explanations is just as opaque as the generation of the predictions. Therefore, our work in this field focuses on better understanding the properties of the generated explanations. We evaluate the downstream performance of a classifier with explanations generated by different model pipelines and compare it to human ratings of the explanations. Our results indicate that the properties that increase the downstream performance differ from those that humans appreciate when evaluating an explanation. Finally, we annotate explanations generated by an LLM for properties that human explanations typically have and discuss the effects those properties have on different user groups. 

While a detailed understanding of the inner workings of LLMs is still unfeasible, I argue that the techniques and analyses presented in this work can help to better understand LLMs, the linguistic knowledge they encode and their decision-making process. Together with knowledge about the models’ architecture, training data and training objective, such techniques can help us develop a robust high-level understanding of LLMs that can guide decisions on their deployment and potential improvements. 

Place, publisher, year, edition, pages
Linköping: Linköping University Electronic Press, 2024. p. 81
Series
Linköping Studies in Science and Technology. Dissertations, ISSN 0345-7524 ; 2364
National Category
Computer and Information Sciences
Identifiers
urn:nbn:se:liu:diva-201985 (URN)10.3384/9789180754712 (DOI)9789180754705 (ISBN)9789180754712 (ISBN)
Public defence
2024-04-18, Ada Lovelace, B-building, Campus Valla, Linköping, 14:00 (English)
Opponent
Supervisors
Available from: 2024-04-02 Created: 2024-04-02 Last updated: 2024-08-01Bibliographically approved
Holmström, O., Kunz, J. & Kuhlmann, M. (2023). Bridging the Resource Gap: Exploring the Efficacy of English and Multilingual LLMs for Swedish. In: Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023): . Paper presented at RESOURCEFUL workshop at NoDaLiDa (pp. 92-110). Tórshavn, the Faroe Islands
Open this publication in new window or tab >>Bridging the Resource Gap: Exploring the Efficacy of English and Multilingual LLMs for Swedish
2023 (English)In: Proceedings of the Second Workshop on Resources and Representations for Under-Resourced Languages and Domains (RESOURCEFUL-2023), Tórshavn, the Faroe Islands, 2023, p. 92-110Conference paper, Published paper (Refereed)
Abstract [en]

Large language models (LLMs) have substantially improved natural language processing (NLP) performance, but training these models from scratch is resource-intensive and challenging for smaller languages. With this paper, we want to initiate a discussion on the necessity of language-specific pre-training of LLMs. We propose how the “one model-many models” conceptual framework for task transfer can be applied to language transfer and explore this approach by evaluating the performance of non-Swedish monolingual and multilingual models’ performance on tasks in Swedish. Our findings demonstrate that LLMs exposed to limited Swedish during training can be highly capable and transfer competencies from English off-the-shelf, including emergent abilities such as mathematical reasoning, while at the same time showing distinct culturally adapted behaviour. Our results suggest that there are resourceful alternatives to language-specific pre-training when creating useful LLMs for small languages.

Place, publisher, year, edition, pages
Tórshavn, the Faroe Islands: , 2023
Keywords
NLP, Natural Language Processing, language model, GPT, monolingual, multilingual, cross-lingual, one model-many models
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-196545 (URN)
Conference
RESOURCEFUL workshop at NoDaLiDa
Funder
CUGS (National Graduate School in Computer Science)
Available from: 2023-08-11 Created: 2023-08-11 Last updated: 2025-02-07
Kunz, J., Jirénius, M., Holmström, O. & Kuhlmann, M. (2022). Human Ratings Do Not Reflect Downstream Utility: A Study of Free-Text Explanations for Model Predictions. In: Jasmijn Bastings, Yonatan Belinkov, Yanai Elazar, Dieuwke Hupkes, Naomi Saphra, Sarah Wiegreffe (Ed.), Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP: . Paper presented at BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, December 8, 2022 (pp. 164-177). Association for Computational Linguistics (ACL), 5, Article ID 2022.blackboxnlp-1.14.
Open this publication in new window or tab >>Human Ratings Do Not Reflect Downstream Utility: A Study of Free-Text Explanations for Model Predictions
2022 (English)In: Proceedings of the Fifth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP / [ed] Jasmijn Bastings, Yonatan Belinkov, Yanai Elazar, Dieuwke Hupkes, Naomi Saphra, Sarah Wiegreffe, Association for Computational Linguistics (ACL) , 2022, Vol. 5, p. 164-177, article id 2022.blackboxnlp-1.14Conference paper, Published paper (Refereed)
Abstract [en]

Models able to generate free-text rationales that explain their output have been proposed as an important step towards interpretable NLP for “reasoning” tasks such as natural language inference and commonsense question answering. However, the relative merits of different architectures and types of rationales are not well understood and hard to measure. In this paper, we contribute two insights to this line of research: First, we find that models trained on gold explanations learn to rely on these but, in the case of the more challenging question answering data set we use, fail when given generated explanations at test time. However, additional fine-tuning on generated explanations teaches the model to distinguish between reliable and unreliable information in explanations. Second, we compare explanations by a generation-only model to those generated by a self-rationalizing model and find that, while the former score higher in terms of validity, factual correctness, and similarity to gold explanations, they are not more useful for downstream classification. We observe that the self-rationalizing model is prone to hallucination, which is punished by most metrics but may add useful context for the classification step.

Place, publisher, year, edition, pages
Association for Computational Linguistics (ACL), 2022
Keywords
Large Language Models, Neural Networks, Transformers, Interpretability, Explainability
National Category
Natural Language Processing Computer Sciences
Identifiers
urn:nbn:se:liu:diva-195615 (URN)10.18653/v1/2022.blackboxnlp-1.14 (DOI)2-s2.0-85152897033 (Scopus ID)
Conference
BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, December 8, 2022
Available from: 2023-06-22 Created: 2023-06-22 Last updated: 2025-02-01Bibliographically approved
Kunz, J. & Kuhlmann, M. (2022). Where Does Linguistic Information Emerge in Neural Language Models?: Measuring Gains and Contributions across Layers. In: Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Paggio, Nianwen Xue, Seokhwan Kim, Younggyun Hahm, Zhong He, Tony Kyungil Lee, Enrico Santus, Francis Bond, Seung-Hoon Na (Ed.), Proceedings of the 29th International Conference on Computational Linguistics: . Paper presented at COLING, October 12–17, 2022 (pp. 4664-4676). Association for Computational Linguistics, 29, Article ID 1.413.
Open this publication in new window or tab >>Where Does Linguistic Information Emerge in Neural Language Models?: Measuring Gains and Contributions across Layers
2022 (English)In: Proceedings of the 29th International Conference on Computational Linguistics / [ed] Nicoletta Calzolari, Chu-Ren Huang, Hansaem Kim, James Pustejovsky, Leo Wanner, Key-Sun Choi, Pum-Mo Ryu, Hsin-Hsi Chen, Lucia Donatelli, Heng Ji, Sadao Kurohashi, Patrizia Paggio, Nianwen Xue, Seokhwan Kim, Younggyun Hahm, Zhong He, Tony Kyungil Lee, Enrico Santus, Francis Bond, Seung-Hoon Na, Association for Computational Linguistics, 2022, Vol. 29, p. 4664-4676, article id 1.413Conference paper, Published paper (Refereed)
Abstract [en]

Probing studies have extensively explored where in neural language models linguistic information is located. The standard approach to interpreting the results of a probing classifier is to focus on the layers whose representations give the highest performance on the probing task. We propose an alternative method that asks where the task-relevant information emerges in the model. Our framework consists of a family of metrics that explicitly model local information gain relative to the previous layer and each layer’s contribution to the model’s overall performance. We apply the new metrics to two pairs of syntactic probing tasks with different degrees of complexity and find that the metrics confirm the expected ordering only for one of the pairs. Our local metrics show a massive dominance of the first layers, indicating that the features that contribute the most to our probing tasks are not as high-level as global metrics suggest.

Place, publisher, year, edition, pages
Association for Computational Linguistics, 2022
Series
Proceedings - International Conference on Computational Linguistics, COLING, ISSN 2951-2093
Keywords
NLP, AI, Language Technology, Computational Linguistics, Machine Learning
National Category
Natural Language Processing
Identifiers
urn:nbn:se:liu:diva-191000 (URN)2-s2.0-85162661424 (Scopus ID)
Conference
COLING, October 12–17, 2022
Available from: 2023-01-12 Created: 2023-01-12 Last updated: 2025-11-17Bibliographically approved
Organisations

Search in DiVA

Show all publications