Shaping Cross-Lingual Transfer: Input Representation as an Inductive Bias

Shaping Cross-Lingual Transfer: Input Representation as an Inductive Bias

Speaker:

‏Benjamin Suter

Date:

2026-05-26 at 15:00–16:00

Location:

Department of Computer Science (Celestijnlaan 200A) -- Perl (room 5.120)

Abstract:

Low-resource languages rely heavily on cross-lingual transfer, yet it remains unclear which properties enable successful transfer across languages. While sharing happens in all layers of a model, the way language is represented at the input level (e.g. through tokenization) acts as an inductive bias, defining the units a model will operate on and shaping what it can immediately recognize as similar or different.
Current tokenizers operate directly on the surface, and indeed, empirical evidence suggests that multilingual transfer is sensitive to spelling and script variation. Ideally, however, what is shared in a model across languages is structural similarity, not surface overlap.
This talk critically assesses implicit assumptions in current tokenization methods, and argues for alternative input representations that are robust to script and spelling variation, balance information density across languages (including previously unseen ones) and emphasize transferable structure. Furthermore, it introduces a dataset for research in low-resource machine translation that was designed to maximize both linguistic and script diversity. This dataset will serve as the basis for future experiments on alternative multilingual input representations.