Cross-domain analysis of discourse markers in European Portuguese

Cabarrão, V.; Moniz, H.; Batista, F.; Ferreira, J.; Trancoso, I.; Mata, A. I.

doi:10.5087/dad.2018.103

Please use this identifier to cite or link to this item: http://hdl.handle.net/10071/16767

Full metadata record

DC Field	Value	Language
dc.contributor.author	Cabarrão, V.	-
dc.contributor.author	Moniz, H.	-
dc.contributor.author	Batista, F.	-
dc.contributor.author	Ferreira, J.	-
dc.contributor.author	Trancoso, I.	-
dc.contributor.author	Mata, A. I.	-
dc.date.accessioned	2018-11-23T13:25:26Z	-
dc.date.available	2018-11-23T13:25:26Z	-
dc.date.issued	2018	-
dc.identifier.issn	2152-9620	-
dc.identifier.uri	http://hdl.handle.net/10071/16767	-
dc.description.abstract	This paper presents an analysis of discourse markers in two spontaneous speech corpora for European Portuguese - university lectures and map-task dialogues - and also in a collection of tweets, aiming at contributing to their categorization, scarcely existent for European Portuguese. Our results show that the selection of discourse markers is domain and speaker dependent. We also found that the most frequent discourse markers are similar in all three corpora, despite tweets containing discourse markers not found in the other two corpora. In this multidisciplinary study, comprising both a linguistic perspective and a computational approach, discourse markers are also automatically discriminated from other structural metadata events, namely sentence-like units and disfluencies. Our results show that discourse markers and disfluencies tend to co-occur in the dialogue corpus, but have a complementary distribution in the university lectures. We used three acoustic-prosodic feature sets and machine learning to automatically distinguish between discourse markers, disfluencies and sentence-like units. Our in-domain experiments achieved an accuracy of about 87% in university lectures and 84% in dialogues, in line with our previous results. The eGeMAPS features, commonly used for other paralinguistic tasks, achieved a considerable performance on our data, especially considering the small size of the feature set. Our results suggest that turn-initial discourse markers are usually easier to classify than disfluencies, a result also previously reported in the literature. We conducted a cross-domain evaluation in order to evaluate the robustness of the models across domains. The results achieved are about 11%-12% lower, but we conclude that data from one domain can still be used to classify the same events in the other. Overall, despite the complexity of this task, these are very encouraging state-of-the-art results. Ultimately, using exclusively acoustic-prosodic cues, discourse markers can be fairly discriminated from disfluencies and SUs. In order to better understand the contribution of each feature, we have also reported the impact of the features in both the dialogues and the university lectures. Pitch features are the most relevant ones for the distinction between discourse markers and disfluencies, namely pitch slopes. These features are in line with the wide pitch range of discourse markers, in a continuum from a very compressed pitch range to a very wide one, expressed by total deaccented material or H+L* L* contours, with upstep H tones.	eng
dc.language.iso	eng	-
dc.publisher	Linguistic Society of America	-
dc.relation	info:eu-repo/grantAgreement/FCT/5876/147282/PT	-
dc.relation	SFRH/BD/96492/2013	-
dc.relation	info:eu-repo/grantAgreement/FCT/SFRH/SFRH%2FBPD%2F95849%2F2013/PT	-
dc.rights	openAccess	-
dc.subject	European Portuguese	eng
dc.subject	Prosody	eng
dc.subject	Speech processing	eng
dc.subject	Structural metadata events	eng
dc.title	Cross-domain analysis of discourse markers in European Portuguese	eng
dc.type	article	-
dc.event.date	2018	-
dc.pagination	79 - 106	-
dc.peerreviewed	yes	-
dc.journal	Dialogue and Discourse	-
dc.volume	9	-
dc.number	1	-
degois.publication.firstPage	79	-
degois.publication.lastPage	106	-
degois.publication.issue	1	-
degois.publication.title	Cross-domain analysis of discourse markers in European Portuguese	eng
dc.date.updated	2019-03-20T17:19:06Z	-
dc.description.version	info:eu-repo/semantics/publishedVersion	-
dc.identifier.doi	10.5087/dad.2018.103	-
dc.subject.fos	Domínio/Área Científica::Ciências Naturais::Ciências da Computação e da Informação	por
dc.subject.fos	Domínio/Área Científica::Engenharia e Tecnologia::Engenharia Eletrotécnica, Eletrónica e Informática	por
dc.subject.fos	Domínio/Área Científica::Humanidades::Línguas e Literaturas	por
iscte.identifier.ciencia	https://ciencia.iscte-iul.pt/id/ci-pub-50912	-
iscte.alternateIdentifiers.scopus	2-s2.0-85049600932	-
Appears in Collections:	CTI-RI - Artigos em revistas científicas internacionais com arbitragem científica

Files in This Item:

File	Description	Size	Format
Cabarrao 2018 - Cross-domain analysis of discourse markers in European.pdf	Versão Editora	2,62 MB	Adobe PDF	View/Open

Show simple item record