Czech Dataset for Cross-lingual Subjectivity Classification

Přibáň, Pavel; Steinberger, Josef

Full metadata record

DC pole	Hodnota	Jazyk
dc.contributor.author	Přibáň, Pavel
dc.contributor.author	Steinberger, Josef
dc.date.accessioned	2023-03-20T11:00:15Z	-
dc.date.available	2023-03-20T11:00:15Z	-
dc.date.issued	2022
dc.identifier.citation	PŘIBÁŇ, P. STEINBERGER, J. Czech Dataset for Cross-lingual Subjectivity Classification. In Proceedings of the Thirteenth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association, 2022. s. 1381-1391. ISBN: 979-10-95546-72-6	cs
dc.identifier.isbn	979-10-95546-72-6
dc.identifier.uri	2-s2.0-85144389838
dc.identifier.uri	http://hdl.handle.net/11025/51716
dc.description.abstract	V tomto článku představujeme nový český dataset pro klasifikaci subjektivity, který obsahuje 10 tisíc manuálně označených subjektivních a objektivních vět z filmových recenzí a popisů filmů. Naší hlavní motivací je poskytnout spolehlivý dataset který může být použit společně s již existujícím anglickým datasetem jako test schopnosti předtrénovaných vícejazyčných modelů pro přenost znalosti mezi češtinou a angličtinou. Dva anotátoři označili dataset a dosáhli 0.83 Cohen Kappa metriky. Dále jsme vytvořili doplňkový dataset který obsahuje 200 tisíc automaticky označených vět. Oba datasety jsou volně k použití pro výzkumné účely. Dále jsme provedli tzv. Fine-tuning pěti předtrénovaných modelů založených na architektuře Transformers pro určení základních výsledků, kde dosahujeme 93.56% úspěšnosti. Dále provádíme experimenty, které mají za cíl ověřit možnosti vícejazyčných modelů pro přenos znalosti mezi jazyky.	cs
dc.format	11 s.	cs
dc.format.mimetype	application/pdf
dc.language.iso	en	en
dc.publisher	European Language Resources Association	en
dc.relation.ispartofseries	EPJ Web of Conferences Volume 269 (2022) EFM19 – Experimental Fluid Mechanics 2019	en
dc.rights	© CC BY-NC-ND 4.0	en
dc.subject	subjektivita	cs
dc.subject	dataset	cs
dc.subject	mezijazyčný	cs
dc.subject	klasifikace	cs
dc.subject	transformers	cs
dc.subject	benchmark	cs
dc.title	Czech Dataset for Cross-lingual Subjectivity Classification	en
dc.title.alternative	Český Dataset pro mezijazyčnou klasifikaci subjektivity	cs
dc.type	konferenční příspěvek	cs
dc.type	ConferenceObject	en
dc.rights.access	openAccess	en
dc.type.version	publishedVersion	en
dc.description.abstract-translated	In this paper, we introduce a new Czech subjectivity dataset of 10k manually annotated subjective and objective sentences from movie reviews and descriptions. Our prime motivation is to provide a reliable dataset that can be used with the existing English dataset as a benchmark to test the ability of pre-trained multilingual models to transfer knowledge between Czech and English and vice versa. Two annotators annotated the dataset reaching 0.83 of the Cohen’s kappa inter-annotator agreement. To the best of our knowledge, this is the first subjectivity dataset for the Czech language. We also created an additional dataset that consists of 200k automatically labeled sentences. Both datasets are freely available for research purposes. Furthermore, we fine-tune five pre-trained BERT-like models to set a monolingual baseline for the new dataset and we achieve 93.56% of accuracy. We fine-tune models on the existing English dataset for which we obtained results that are on par with the current state-of-the-art results. Finally, we perform zero-shot cross-lingual subjectivity classification between Czech and English to verify the usability of our dataset as the cross-lingual benchmark. We compare and discuss the cross-lingual and monolingual results and the ability of multilingual models to transfer knowledge between languages.	en
dc.subject.translated	subjectivity	en
dc.subject.translated	dataset	en
dc.subject.translated	Czech	en
dc.subject.translated	cross-lingua	en
dc.subject.translated	classification	en
dc.subject.translated	transformers	en
dc.subject.translated	benchmark	en
dc.type.status	Peer-reviewed	en
dc.identifier.document-number	889371701052
dc.identifier.obd	43936946
dc.project.ID	SGS-2022-016/Pokročilé metody zpracování a analýzy dat	cs
Vyskytuje se v kolekcích:	Konferenční příspěvky / Conference Papers (KIV) OBD

Soubory připojené k záznamu:

Soubor	Velikost	Formát
Přibáň, Steinberger paper-LREC.pdf	347,57 kB	Adobe PDF	Zobrazit/otevřít

Zobrazit minimální záznam Zobrazit statistiky

Použijte tento identifikátor k citaci nebo jako odkaz na tento záznam: http://hdl.handle.net/11025/51716

Všechny záznamy v DSpace jsou chráněny autorskými právy, všechna práva vyhrazena.

hledání

navigace