Transliteration in Any Language with Surrogate Languages

09/14/2016
by   Stephen Mayhew, et al.
0

We introduce a method for transliteration generation that can produce transliterations in every language. Where previous results are only as multilingual as Wikipedia, we show how to use training data from Wikipedia as surrogate training for any language. Thus, the problem becomes one of ranking Wikipedia languages in order of suitability with respect to a target language. We introduce several task-specific methods for ranking languages, and show that our approach is comparable to the oracle ceiling, and even outperforms it in some cases.

READ FULL TEXT

Please sign up or login with your details

Forgot password? Click here to reset