Morzsák

Oldal címe

Evaluating Latent Semantic Pre-training for Fine-grained Word Sense Disambiguation

Címlapos tartalom

Understanding word meaning in context is a central problem in Natural Language Processing, most directly addressed by word sense disambiguation (WSD). While modern contextual language models have improved WSD performance, it remains unclear how the choice of pre-training objective influences the semantic representations used in different WSD approaches. In this paper, we investigate masked latent semantic modeling (MLSM) as a semantic pre-training objective for WSD, focusing on its behavior across diverse WSD approaches. We contrast the use of MLSM pre-trained models to that of traditionally pre-trained ones for a wide range of WSD modeling scenarios, including centroid and sparse representation-based probing methods, as well as supervised disambiguation systems. Experimental results on standard WSD benchmarks show that relying on MLSM yields consistent improvements in disambiguation performance across models and datasets.