Events
Shaping Cross-Lingual Transfer: Input Representation as an Inductive Bias
Speaker:
Benjamin SuterDate:
2026-05-26 at 15:00–16:00Location:
Department of Computer Science (Celestijnlaan 200A) -- Perl (room 5.120)Abstract:
Low-resource languages rely heavily on cross-lingual transfer, yet it remains unclear which properties enable successful transfer across languages. While sharing happens in all layers of a model, the way language is represented at the input level (e.g. through tokenization) acts as an inductive bias, defining the units a model will operate on and shaping what it can immediately recognize as similar or different.Current tokenizers operate directly on the surface, and indeed, empirical evidence suggests that multilingual transfer is sensitive to spelling and script variation. Ideally, however, what is shared in a model across languages is structural similarity, not surface overlap.
This talk critically assesses implicit assumptions in current tokenization methods, and argues for alternative input representations that are robust to script and spelling variation, balance information density across languages (including previously unseen ones) and emphasize transferable structure. Furthermore, it introduces a dataset for research in low-resource machine translation that was designed to maximize both linguistic and script diversity. This dataset will serve as the basis for future experiments on alternative multilingual input representations.
Towards Idea Discovery in Buddhist Corpora: Extracting and Matching Ideas from Corpora via Idiomatic Expressions
Speaker:
Kai Golan HashiloniDate:
2026-02-03 at 15:00–16:00Location:
Department of Computer Science (Celestijnlaan 200A) -- Java (room 5.152)Abstract:
Understanding meaning beyond surface form remains a central challenge for large language models (LLMs). In this talk, I present a trajectory of recent and ongoing research on how LLMs engage with nonliteral and context-dependent meaning, focusing on idiomatic expressions as a test case. This work represents a step toward the broader goal of Idea Discovery; enabling systems to extract and cluster ideas or concepts from large text corpora in an unsupervised manner. These approaches should then be applied in my research for the Intellexus Project, where our goal is to design and integrate NLP-based solutions for the analysis of Buddhist corpora, with a focus on Sanskrit and Tibetan. I show that LLMs can identify idioms in context, even across languages and with minimal task-specific supervision. However, this apparent success is fragile. Models struggle to reliably link idiomatic expressions to semantically equivalent literal paraphrases. Together, these findings highlight both the promise and the limitations of LLMs as models of conceptual meaning, and they expose a persistent gap between current systems and human-level interpretation.
Bio:
Kai is a Computer Science PhD student at Reichman University, Israel, where he also earned his BSc in Computer Science and Entrepreneurship and his MSc in Computer Science. This included, interestingly, 2.5 years of exchange period in TU Darmstadt, Germany, where he also learned German (the hard way).He focuses on natural language processing (NLP) and is a lead researcher at the Intellexus Project, where he investigates the design and integration of NLP-based solutions for the study of ancient Buddhist corpora in Sanskrit and Tibetan. His work involves training large language models (LLMs), designing and curating downstream tasks and evaluation benchmarks, LLMs' explainability studies, and more.
Kai's research interests lie in multilingual NLP and low-resource languages, and in addressing linguistic and semantic challenges, such as idiom processing, using LLMs. → Website
Language Similarity in Machine Translation: A Typological Perspective
Speaker:
drs. Esther PloegerDate:
2024-08-01 at 13:00–14:00Location:
Department of Computer Science (Celestijnenlaan 200A) -- Java (room 5.152)Abstract:
State-of-the-art performance in machine translation is currently achieved by training models on an ever larger amount of text. While this data-driven approach has increased translation quality substantially in some cases, languages for which little such data is available often remain underserved. Additionally, this approach is increasingly computationally expensive. Linguistic typology provides potential for mitigating these issues. Typologists have systematically analysed the similarities and differences between many of the world’s languages, resulting in comprehensive linguistic descriptions and databases. This talk will address challenges and opportunities in applying what we know about languages to make translation models more efficient and accurate in low-resource settings.
Bio:
Esther Ploeger is a PhD student at the Department of Computer Science at Aalborg University, Copenhagen. In her research, she focuses on leveraging knowledge about language and cross-linguistic tendencies (linguistic typology) in practical natural language processing applications, such as machine translation. Prior to this, she obtained a BSc. And MSc. in Information Science at the Univeristy of Groningen in The Netherlands. → WebsiteWhen Language Models Meet Words
Speaker:
Dr. Yuval PinterSenior Lecturer at the Department of Computer Science of Ben-Gurion University