Skip to content

About

Dataset for TALLIP2019 paper "Ancient-Modern Chinese Translation with a New Large Training Dataset"

Resources

Stars

0 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

Ancient-Modern Chinese Translation with a New Large Training Dataset

This repo contains the dataset built in the following paper:

Ancient-Modern Chinese Translation with a New Large Training Dataset. Dayiheng Liu, Kexin Yang, Qian Qu, Jiancheng Lv, TALLIP 2019 [arXiv]

Overview

We create a new large-scale Ancient-Modern Chinese parallel corpus which contains 1.24M bilingual pairs. To our best knowledge, this is the first large high-quality Ancient-Modern Chinese dataset.

Dataset

We plan to gradually release the dataset. The way to get the data is as follows:

Please scan the completed agreement and send it to losinuris[AT]gmail.com. A notification email will be sent to the email address and downloading of the Dataset will be authorized once the procedure is approved.

About

Dataset for TALLIP2019 paper "Ancient-Modern Chinese Translation with a New Large Training Dataset"

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors