Investigating the Structural Properties of Linguistic Biases in Multilingual Language Models


Mantri R., Chen S., Wang Y., ATAMAN D.

Information (Switzerland), vol.17, no.5, 2026 (ESCI, Scopus)

  • Publication Type: Article / Article
  • Volume: 17 Issue: 5
  • Publication Date: 2026
  • Doi Number: 10.3390/info17050498
  • Journal Name: Information (Switzerland)
  • Journal Indexes: Emerging Sources Citation Index (ESCI), Scopus, Aerospace Database, Compendex, INSPEC, Library, Information Science & Technology Abstracts (LISTA), Directory of Open Access Journals, Information Science & Technology Abstracts (LISTA), Academic Search Ultimate (EBSCO), Technology Collection (ProQuest)
  • Keywords: generalization, large language models, learning bias, multilinguality, representation learning
  • Middle East Technical University Affiliated: Yes

Abstract

As large language models (LLMs) scale to cover more languages, their potential to support low-resource settings becomes increasingly promising. However, the mechanisms underlying cross-lingual transfer and the factors that facilitate it remain insufficiently understood. Prior work has highlighted the role of linguistic similarity—particularly syntactic structure—in enabling transfer across languages. In this study, we present a broad empirical analysis of how multilingual LLMs encode and relate structural information across languages with varying typological properties. We combine multiple complementary methods, including hidden-state similarity analysis, typological correlation, probing for syntactic features, and attention-based structural comparisons, across four multilingual models and thirteen languages. Our findings show consistent correlations between representational similarity and syntactic relatedness, suggesting that structural properties of language influence how information is organized and shared across languages. We further observe that attention-derived structures exhibit partial alignment with gold-standard syntax, though this alignment should be interpreted as heuristic rather than direct evidence of syntactic encoding. Overall, our results provide a comparative empirical perspective on cross-lingual structural bias in multilingual LLMs and highlight the importance of careful methodological interpretation when linking representation geometry to linguistic structure.