Apple Workshop on Machine Learning for Health: Pre-trained Model Representations and their Robustness against Noise for Speech Emotion Analysis
AuthorsVikram Mitra (Apple)
Apple Workshop on Machine Learning for Health: Pre-trained Model Representations and their Robustness against Noise for Speech Emotion Analysis
AuthorsVikram Mitra (Apple)
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions
August 20, 2026research area Speech and Natural Language Processing
Cross-lingual knowledge transfer is critical for building high-performing multilingual language models for languages with insufficient training data. When target language data is scarce, the knowledge required for many downstream tasks involving scientific reasoning, commonsense inference, and world knowledge must be acquired primarily from the high-resource language, making effective knowledge transfer essential. Existing methods for improving…
Scaling Laws for Mixture Pretraining Under Data Constraints
August 20, 2026research area Methods and Algorithms, research area Speech and Natural Language Processing
As language models scale, the amount of data they require grows – yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same…