Learn to Code-Switch: Data Augmentation using Copy Mechanism on Language Modeling

10/24/2018
by   Genta Indra Winata, et al.
0

Building large-scale datasets for training code-switching language models is challenging and very expensive. To alleviate this problem parallel corpus has been a major workaround. However, existing solutions use linguistic constraints which may not capture the real data distribution. In this work, we propose a novel method for learning how to generate code-switching sentences from parallel corpora. Our model uses a Seq2Seq in combination with pointer networks to align and choose words from the monolingual sentences and form a grammatical code-switching sentence. In our experiment, we show that by training a language model using the generated sentences improve the perplexity score by around 10

READ FULL TEXT

page 1

page 2

page 3

page 4

research
09/18/2019

Code-Switched Language Models Using Neural Based Synthetic Data from Parallel Sentences

Training code-switched language models is difficult due to lack of data ...
research
11/06/2018

Code-switching Sentence Generation by Generative Adversarial Networks and its Application to Data Augmentation

Code-switching is about dealing with alternative languages in speech or ...
research
04/13/2021

Multilingual Transfer Learning for Code-Switched Language and Speech Neural Modeling

In this thesis, we address the data scarcity and limitations of linguist...
research
12/12/2021

Improving Code-switching Language Modeling with Artificially Generated Texts using Cycle-consistent Adversarial Networks

This paper presents our latest effort on improving Code-switching langua...
research
07/14/2021

From Machine Translation to Code-Switching: Generating High-Quality Code-Switched Text

Generating code-switched text is a problem of growing interest, especial...
research
12/23/2020

Code Switching Language Model Using Monolingual Training Data

Training a code-switching (CS) language model using only monolingual dat...
research
05/30/2018

Code-Switching Language Modeling using Syntax-Aware Multi-Task Learning

Lack of text data has been the major issue on code-switching language mo...

Please sign up or login with your details

Forgot password? Click here to reset