Header

Suche

Text Technology/Digital Linguistics Colloquium HS 2026

Time & Location: every 2 weeks on Tuesdays 10:15-12:00 in room BIM 4-033.

Online participation via the MS Teams Team CL Colloquium is also possible.

Responsible: Marius Huber

Colloquium Schedule

 
15.09.2026 Mark Fišel Jana Massoud & Sina Ahmadi
29.09.2026 Yiyang Chen Tanja Samardžić
13.10.2026 Anna Bondar Michelle Wastl
27.10.2026 Cui Ding Jan Brasser
10.11.2026 Bastian Rieck

Jason Armitage

24.11.2026 Kirill Semenov Andrianos Michail
08.12.2026 Longqian Ming Gerard Sant Muniesa

15 Sep 2026

Mark Fišel: Modelling Extremely Low-resource Finno-Ugric Languages and Dialects

I will describe an on-going effort on developing NLP tools for the Finno-Ugric family, the members of which range from languages with 1M+ speakers and "an army and fleet" to critically endangered dialects and languages with few resources and even fewer speakers. The latest development direction in this work is on multi-tasking models and general-purpose LMs, which is, however, hindered by the availability of data, low domain coverage, diverse quality in data from different sources and variety in orthography. I will present the already developed models and performed experiments, and will be happy to discuss open questions.

Jana Massoud & Sina Ahmadi: Towards a quantitative language resource index

The term "low-resource" in NLP is too vague to be useful. This talk presents our project that aims to move towards a quantitative resource index that scores 8,168 languages across six key dimensions (from corpora and models to dictionaries). Our analysis reveals a heavily skewed landscape where 75% of the world's languages are entirely absent from contemporary NLP. We validate this index against downstream tasks, offering a new, fine-grained tool to accurately situate languages in the tech ecosystem and identify where resources are needed most.

29 Sep 2026

Yiyang Chen: Towards Automatic, Label-Free Expressive TTS

 How can a text-to-speech (TTS) system infer an appropriate emotional expression from text and use it to guide speech synthesis? This talk explores automatic emotion prediction and control in zero-shot TTS. I will introduce our progression from selecting emotional speech references using predicted valence, arousal, and dominance (VAD) to EmoPilot, which predicts emotion cues directly in the TTS backend’s learned emotion space. These cues support reference-audio retrieval or direct conditioning, bypassing explicit emotion-category and VAD predictions. I will then discuss our next steps towards context-aware emotion prediction: the same sentence can convey different emotions depending on its surrounding context. Our upcoming Romansh long-form audio generation project provides a motivating setting, with training speech limited in both quantity and emotional expressiveness. Can emotional expression learned from speech in other languages transfer to Romansh synthesis? Reference and synthesized audio examples will illustrate the research questions and provide a starting point for discussing future directions and applications.

Tanja Samardžić: Controlled token allocation in optimising multilingual vocabularies

Multilingual (pre-)training of large language models is necessary but it is prone to various biases due to the uneven data availability and structural differences between languages. This problem becomes apparent already at the level of tokenisation: if no care is taken, a multilingual vocabulary (constructed by means of BPE) ends up dominated by items representing high-resource languages. To prevent this, various controlling techniques have been proposed, from upsampling tokens from low-resource languages to minimising token sharing across languages and setting cross-lingual parity as one of the BPE's objectives. In this talk, I will present experiments testing potential advantages of several new techniques aimed at designing better multilingual vocabularies. Our preliminary findings suggest that controlled token allocation is more beneficial with larger vocabularies allowing systematic improvements in model-intrinsic metrics (bits-per-byte, information parity) and also in some downstream tasks (belebele).

13 Oct 2026

Anna Bondar

TBA

Michelle Wastl

TBA

27 Oct 2026

Cui Ding

TBA

Jan Brasser

TBA 

10 Nov 2026

Bastian Rieck

TBA

Jason Armitage

TBA

24 Nov 2026

Kirill Semenov

TBA

Andrianos Michail

TBA

8 Dec 2026

Longqian Ming

TBA

Gerard Sant Muniesa

TBA