ResearchTrend.AI
  • Papers
  • Communities
  • Events
  • Blog
  • Pricing
Papers
Communities
Social Events
Terms and Conditions
Pricing
Parameter LabParameter LabTwitterGitHubLinkedInBlueskyYoutube

© 2025 ResearchTrend.AI, All rights reserved.

  1. Home
  2. Papers
  3. 2505.24561
15
1

Improving Language and Modality Transfer in Translation by Character-level Modeling

30 May 2025
Ioannis Tsiamas
David Dale
Marta R. Costa-jussá
ArXiv (abs)PDFHTML
Main:8 Pages
3 Figures
Bibliography:4 Pages
17 Tables
Appendix:5 Pages
Abstract

Current translation systems, despite being highly multilingual, cover only 5% of the world's languages. Expanding language coverage to the long-tail of low-resource languages requires data-efficient methods that rely on cross-lingual and cross-modal knowledge transfer. To this end, we propose a character-based approach to improve adaptability to new languages and modalities. Our method leverages SONAR, a multilingual fixed-size embedding space with different modules for encoding and decoding. We use a teacher-student approach with parallel translation data to obtain a character-level encoder. Then, using ASR data, we train a lightweight adapter to connect a massively multilingual CTC ASR model (MMS), to the character-level encoder, potentially enabling speech translation from 1,000+ languages. Experimental results in text translation for 75 languages on FLORES+ demonstrate that our character-based approach can achieve better language transfer than traditional subword-based models, especially outperforming them in low-resource settings, and demonstrating better zero-shot generalizability to unseen languages. Our speech adaptation, maximizing knowledge transfer from the text modality, achieves state-of-the-art results in speech-to-text translation on the FLEURS benchmark on 33 languages, surpassing previous supervised and cascade models, albeit being a zero-shot model with minimal supervision from ASR data.

View on arXiv
@article{tsiamas2025_2505.24561,
  title={ Improving Language and Modality Transfer in Translation by Character-level Modeling },
  author={ Ioannis Tsiamas and David Dale and Marta R. Costa-jussà },
  journal={arXiv preprint arXiv:2505.24561},
  year={ 2025 }
}
Comments on this paper