Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Turkish NLP relies on aging BERT-era models from 2020 or massively multilingual models that under-represent Turkish. Only one novel publicly available Turkish encoder leverages the architectural advances introduced in ModernBERT (RoPE, GLU activations, alternating local-global attention, unpadded training). We introduce ModernBERT-TR, a 150M-parameter encoder pretrained from scratch on 144.4 billion Turkish tokens using the ModernBERT architecture. We train a custom 50K tokenizer optimized for Turkish morphology, curate a two-source balanced training corpus, and conduct systematic hyperparameter ablations. Under a frozen-encoder linear-probing evaluation protocol, ModernBERT-TR achieves a 60.2% average score across 11 diverse Turkish NLP tasks, surpassing the next-best model by 13.1% relative and outperforming the conventional BERTurk baseline by 70.3% relative. Under full fine-tuning on the 28-task TabiBench benchmark, ModernBERT-TR scores 77.28, matching the concurrent TabiBERT within 0.30 points and leading in 5 of 8 task categories despite training on 7× fewer tokens. These gains are achieved with 623 GPUhours on 4×NVIDIA H100 GPUs. We release the model weights, tokenizer, evaluation code, and training configuration to accelerate Turkish NLP research.</p>

Show More

Keywords

turkish training modernberttr models from

Related Articles

PORE

About

Connect