Please use this identifier to cite or link to this item:
http://hdl.handle.net/1893/37466| Appears in Collections: | Computing Science and Mathematics Journal Articles |
| Peer Review Status: | Refereed |
| Title: | Building a Turkish UCCA dataset |
| Author(s): | Bölücü, Necva Can, Burcu |
| Contact Email: | burcu.can@stir.ac.uk |
| Keywords: | Universal Conceptual Cognitive Annotation UCCA Semantic representation METU-Sabanci Turkish Treebank dataset |
| Issue Date: | Jan-2025 |
| Date Deposited: | 8-Oct-2025 |
| Citation: | Bölücü N & Can B (2025) Building a Turkish UCCA dataset. Can Buglalilar B (Supervisor) <i>Natural Language Processing</i>, 31 (1), pp. 111-149. https://doi.org/10.1017/nlp.2024.36 |
| Abstract: | it to a logical form that can be processed and understood by machines. It is utilised by many applications in natural language processing (NLP), particularly in tasks relevant to natural language understanding(NLU). Due to the widespread use of semantic parsing in NLP, many semantic representation schemes with different forms have been proposed; Universal Conceptual Cognitive Annotation (UCCA) is one of them. UCCA is a cross-lingual semantic annotation framework that allows easy annotation without requiring substantial linguistic knowledge. UCCA-annotated datasets have been released so far for English, French, German, Russian, and Hebrew. In this paper, we present a UCCA-annotated Turkish dataset of 400 sentences that are obtained from the METU-Sabanci Turkish Treebank. We provide the UCCA annotation specifications defined for the Turkish language so that it can be extended further. We followed a semiautomatic annotation approach, where an external semantic parser is utilised for the initial annotation of the dataset, which is manually revised by two annotators. We used the same semantic parser model to evaluate the dataset with zero-shot and few-shot learning, demonstrating that even a small sample set from the target language in the training data has a notable impact on the performance of the parser (15.6% and 2.5% gain over zero-shot for labelled and unlabelled results, respectively). |
| DOI Link: | 10.1017/nlp.2024.36 |
| Rights: | C The Author(s), 2024. Published by Cambridge University Press. This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted re-use, distribution and reproduction, provided the original article is properly cited. |
| Licence URL(s): | http://creativecommons.org/licenses/by/4.0/ |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| Advancing Inclusive Brain Health and Dementia Care for People with Intellectual and.pdf | Fulltext - Published Version | 13.63 MB | Adobe PDF | View/Open |
This item is protected by original copyright |
A file in this item is licensed under a Creative Commons License
Items in the Repository are protected by copyright, with all rights reserved, unless otherwise indicated.
The metadata of the records in the Repository are available under the CC0 public domain dedication: No Rights Reserved https://creativecommons.org/publicdomain/zero/1.0/
If you believe that any material held in STORRE infringes copyright, please contact library@stir.ac.uk providing details and we will remove the Work from public display in STORRE and investigate your claim.
