Text Technology/Digital Linguistics Colloquium HS 2026
Time & Location: every 2 weeks on Tuesdays 10:15-12:00 in room BIM 4-033.
Online participation via the MS Teams Team CL Colloquium is also possible.
Responsible: Marius Huber
Colloquium Schedule
| 15.09.2026 | Mark Fišel | Jana Massoud & Sina Ahmadi |
| 29.09.2026 | Yiyang Chen | Tanja Samardžić |
| 13.10.2026 | Anna Bondar | Michelle Wastl |
| 27.10.2026 | Cui Ding | Jan Brasser |
| 10.11.2026 | Bastian Rieck |
Jason Armitage |
| 24.11.2026 | Kirill Semenov | Andrianos Michail |
| 08.12.2026 | Longqian Ming | Gerard Sant Muniesa |
15 Sep 2026
Mark Fišel: Modelling Extremely Low-resource Finno-Ugric Languages and Dialects
I will describe an on-going effort on developing NLP tools for the Finno-Ugric family, the members of which range from languages with 1M+ speakers and "an army and fleet" to critically endangered dialects and languages with few resources and even fewer speakers. The latest development direction in this work is on multi-tasking models and general-purpose LMs, which is, however, hindered by the availability of data, low domain coverage, diverse quality in data from different sources and variety in orthography. I will present the already developed models and performed experiments, and will be happy to discuss open questions.
Jana Massoud & Sina Ahmadi: Towards a quantitative language resource index
The term "low-resource" in NLP is too vague to be useful. This talk presents our project that aims to move towards a quantitative resource index that scores 8,168 languages across six key dimensions (from corpora and models to dictionaries). Our analysis reveals a heavily skewed landscape where 75% of the world's languages are entirely absent from contemporary NLP. We validate this index against downstream tasks, offering a new, fine-grained tool to accurately situate languages in the tech ecosystem and identify where resources are needed most.
29 Sep 2026
Yiyang Chen: Towards Automatic, Label-Free Expressive TTS
How can a text-to-speech (TTS) system infer an appropriate emotional expression from text and use it to guide speech synthesis? This talk explores automatic emotion prediction and control in zero-shot TTS. I will introduce our progression from selecting emotional speech references using predicted valence, arousal, and dominance (VAD) to EmoPilot, which predicts emotion cues directly in the TTS backend’s learned emotion space. These cues support reference-audio retrieval or direct conditioning, bypassing explicit emotion-category and VAD predictions. I will then discuss our next steps towards context-aware emotion prediction: the same sentence can convey different emotions depending on its surrounding context. Our upcoming Romansh long-form audio generation project provides a motivating setting, with training speech limited in both quantity and emotional expressiveness. Can emotional expression learned from speech in other languages transfer to Romansh synthesis? Reference and synthesized audio examples will illustrate the research questions and provide a starting point for discussing future directions and applications.
Tanja Samardžić: Controlled token allocation in optimising multilingual vocabularies
Multilingual (pre-)training of large language models is necessary but it is prone to various biases due to the uneven data availability and structural differences between languages. This problem becomes apparent already at the level of tokenisation: if no care is taken, a multilingual vocabulary (constructed by means of BPE) ends up dominated by items representing high-resource languages. To prevent this, various controlling techniques have been proposed, from upsampling tokens from low-resource languages to minimising token sharing across languages and setting cross-lingual parity as one of the BPE's objectives. In this talk, I will present experiments testing potential advantages of several new techniques aimed at designing better multilingual vocabularies. Our preliminary findings suggest that controlled token allocation is more beneficial with larger vocabularies allowing systematic improvements in model-intrinsic metrics (bits-per-byte, information parity) and also in some downstream tasks (belebele).
13 Oct 2026
TBA
TBA
27 Oct 2026
TBA
TBA
10 Nov 2026
TBA
TBA
24 Nov 2026
TBA
TBA
8 Dec 2026
TBA
TBA