Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean

Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean. Oh, H., Han, J. Y., Choe, H., Park, S., He, H., Choi, J. D., Han, N., Hwang, J. D., & Kim, H. In Proceedings of the International Conference on Parsing Technologies, of IWPT'20, pages 122–131, 2020.

Paper

Paper abstract bibtex

In this paper, we first open on important issues regarding the Penn Korean Universal Treebank (PKT-UD) and address these issues by revising the entire corpus manually with the aim of producing cleaner UD annotations that are more faithful to Korean grammar. For compatibility to the rest of UD corpora, we follow the UDv2 guidelines, and extensively revise the part-of-speech tags and the dependency relations to reflect morphological features and flexible word-order aspects in Korean. The original and the revised versions of PKT-UD are experimented with transformer-based parsing models using biaffine attention. The parsing model trained on the revised corpus shows a significant improvement of 3.0% in labeled attachment score over the model trained on the previous corpus. Our error analysis demonstrates that this revision allows the parsing model to learn relations more robustly, reducing several critical errors that used to be made by the previous model.

@inproceedings{oh:20a,
	abstract = {In this paper, we first open on important issues regarding the Penn Korean Universal Treebank (PKT-UD) and address these issues by revising the entire corpus manually with the aim of producing cleaner UD annotations that are more faithful to Korean grammar. For compatibility to the rest of UD corpora, we follow the UDv2 guidelines, and extensively revise the part-of-speech tags and the dependency relations to reflect morphological features and flexible word-order aspects in Korean. The original and the revised versions of PKT-UD are experimented with transformer-based parsing models using biaffine attention. The parsing model trained on the revised corpus shows a significant improvement of 3.0% in labeled attachment score over the model trained on the previous corpus. Our error analysis demonstrates that this revision allows the parsing model to learn relations more robustly, reducing several critical errors that used to be made by the previous model.},
	author = {Oh, Hwan and Han, Ji Yoon and Choe, Hyonsu and Park, Seokwon and He, Han and Choi, Jinho D. and Han, Na-Rae and Hwang, Jena D. and Kim, Hansaem},
	booktitle = {Proceedings of the International Conference on Parsing Technologies},
	date-added = {2020-05-19 14:19:28 -0400},
	date-modified = {2020-07-20 18:25:43 -0400},
	keywords = {emorynlp; selected},
	pages = {122--131},
	series = {IWPT'20},
	title = {{Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean}},
	url = {https://iwpt20.sigparse.org},
	url_paper = {https://www.aclweb.org/anthology/2020.iwpt-1.13},
	year = {2020},
	Bdsk-Url-1 = {https://iwpt20.sigparse.org}}

Downloads: 0

{"_id":"KQPGjyjiquCogfows","bibbaseid":"oh-han-choe-park-he-choi-han-hwang-etal-analysisofthepennkoreanuniversaldependencytreebankpktudmanualrevisiontobuildrobustparsingmodelinkorean-2020","authorIDs":["4T4rWesgALz7HEKzR","6z2XxaoKxGSwagg33","8raaF5fbapBh6qFdj","BF8vBbjNNnXrDzvrj","CcpeJM53kFxMEvGwX","Gm5wHcR7N8fvkrTuQ","HnTR3Q9CXNgnPhoxg","JZLkTx6GcyGsgwHX3","LMoqLgq5XcAihoPzq","MYj7kDncHFkXGeeCB","PSk88DNdiqd6eoxkw","R3rsXoZQyJyzZrqez","Sunn7D3jhr4Pcn3MQ","T3oZ2xTrP2GWGQDQc","Tuy9dKJ6u225eZr5G","abWTeKWpsYJT72gTK","azJypGPEA4wmFzsbu","bisY5ZvhXb8JZ5XiZ","bjopqAS2SEAqoCxTn","cayeKFYeNGtfJFezG","e2vvcap4BtnvAvoke","fjgwCB85tXjNXyEzZ","fooGikSJyfmCEHJuD","iv47sy5Ly34dP8mJj","mPN22sM2Qh5k9588D","n5GLezpmstkXE28LZ","oSne76dP3vzoas8HZ","pb6HWJfm9rdvQNhjz","rEtLdWCDtG8gkC3Sw","v3n6QN8LAJ3Te6ZNM","z6PPBZDvm345wyXbd"],"author_short":["Oh, H.","Han, J. Y.","Choe, H.","Park, S.","He, H.","Choi, J. D.","Han, N.","Hwang, J. D.","Kim, H."],"bibdata":{"bibtype":"inproceedings","type":"inproceedings","abstract":"In this paper, we first open on important issues regarding the Penn Korean Universal Treebank (PKT-UD) and address these issues by revising the entire corpus manually with the aim of producing cleaner UD annotations that are more faithful to Korean grammar. For compatibility to the rest of UD corpora, we follow the UDv2 guidelines, and extensively revise the part-of-speech tags and the dependency relations to reflect morphological features and flexible word-order aspects in Korean. The original and the revised versions of PKT-UD are experimented with transformer-based parsing models using biaffine attention. The parsing model trained on the revised corpus shows a significant improvement of 3.0% in labeled attachment score over the model trained on the previous corpus. Our error analysis demonstrates that this revision allows the parsing model to learn relations more robustly, reducing several critical errors that used to be made by the previous model.","author":[{"propositions":[],"lastnames":["Oh"],"firstnames":["Hwan"],"suffixes":[]},{"propositions":[],"lastnames":["Han"],"firstnames":["Ji","Yoon"],"suffixes":[]},{"propositions":[],"lastnames":["Choe"],"firstnames":["Hyonsu"],"suffixes":[]},{"propositions":[],"lastnames":["Park"],"firstnames":["Seokwon"],"suffixes":[]},{"propositions":[],"lastnames":["He"],"firstnames":["Han"],"suffixes":[]},{"propositions":[],"lastnames":["Choi"],"firstnames":["Jinho","D."],"suffixes":[]},{"propositions":[],"lastnames":["Han"],"firstnames":["Na-Rae"],"suffixes":[]},{"propositions":[],"lastnames":["Hwang"],"firstnames":["Jena","D."],"suffixes":[]},{"propositions":[],"lastnames":["Kim"],"firstnames":["Hansaem"],"suffixes":[]}],"booktitle":"Proceedings of the International Conference on Parsing Technologies","date-added":"2020-05-19 14:19:28 -0400","date-modified":"2020-07-20 18:25:43 -0400","keywords":"emorynlp; selected","pages":"122–131","series":"IWPT'20","title":"Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean","url":"https://iwpt20.sigparse.org","url_paper":"https://www.aclweb.org/anthology/2020.iwpt-1.13","year":"2020","bdsk-url-1":"https://iwpt20.sigparse.org","bibtex":"@inproceedings{oh:20a,\n\tabstract = {In this paper, we first open on important issues regarding the Penn Korean Universal Treebank (PKT-UD) and address these issues by revising the entire corpus manually with the aim of producing cleaner UD annotations that are more faithful to Korean grammar. For compatibility to the rest of UD corpora, we follow the UDv2 guidelines, and extensively revise the part-of-speech tags and the dependency relations to reflect morphological features and flexible word-order aspects in Korean. The original and the revised versions of PKT-UD are experimented with transformer-based parsing models using biaffine attention. The parsing model trained on the revised corpus shows a significant improvement of 3.0% in labeled attachment score over the model trained on the previous corpus. Our error analysis demonstrates that this revision allows the parsing model to learn relations more robustly, reducing several critical errors that used to be made by the previous model.},\n\tauthor = {Oh, Hwan and Han, Ji Yoon and Choe, Hyonsu and Park, Seokwon and He, Han and Choi, Jinho D. and Han, Na-Rae and Hwang, Jena D. and Kim, Hansaem},\n\tbooktitle = {Proceedings of the International Conference on Parsing Technologies},\n\tdate-added = {2020-05-19 14:19:28 -0400},\n\tdate-modified = {2020-07-20 18:25:43 -0400},\n\tkeywords = {emorynlp; selected},\n\tpages = {122--131},\n\tseries = {IWPT'20},\n\ttitle = {{Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean}},\n\turl = {https://iwpt20.sigparse.org},\n\turl_paper = {https://www.aclweb.org/anthology/2020.iwpt-1.13},\n\tyear = {2020},\n\tBdsk-Url-1 = {https://iwpt20.sigparse.org}}\n\n","author_short":["Oh, H.","Han, J. Y.","Choe, H.","Park, S.","He, H.","Choi, J. D.","Han, N.","Hwang, J. D.","Kim, H."],"key":"oh:20a","id":"oh:20a","bibbaseid":"oh-han-choe-park-he-choi-han-hwang-etal-analysisofthepennkoreanuniversaldependencytreebankpktudmanualrevisiontobuildrobustparsingmodelinkorean-2020","role":"author","urls":{"Paper":"https://iwpt20.sigparse.org"," paper":"https://www.aclweb.org/anthology/2020.iwpt-1.13"},"keyword":["emorynlp; selected"],"metadata":{"authorlinks":{"choi, j":"http://www.cs.emory.edu/"}},"html":""},"bibtype":"inproceedings","biburl":"http://www.mathcs.emory.edu/~choi/cv/jinho_choi-20210601.bib","creationDate":"2020-05-23T03:23:45.219Z","downloads":6,"keywords":["emorynlp; selected"],"search_terms":["analysis","penn","korean","universal","dependency","treebank","pkt","manual","revision","build","robust","parsing","model","korean","oh","han","choe","park","he","choi","han","hwang","kim"],"title":"Analysis of the Penn Korean Universal Dependency Treebank (PKT-UD): Manual Revision to Build Robust Parsing Model in Korean","year":2020,"dataSources":["WRbWtphN7JFZaJSwS","KCe4LtCfLaE5R9apZ"]}