Ciberdúvidas da Língua Portuguesa: European Portuguese Linguistics Paraphrase Benchmark - 1000-Question Scaled Evaluation Set (5000 Queries)
| dc.contributor.author | Moura, P. | |
| dc.contributor.author | Batista, F. | |
| dc.contributor.author | Lopes, A. | |
| dc.date.accessioned | 2026-09-29T14:35:33Z | |
| dc.date.issued | 2026-09-17 | |
| dc.description.abstract | This dataset contains a large scale paraphrase-based query benchmark built from 1000 original questions randomly sampled from Ciberdúvidas da Língua Portuguesa, an expert-curated Portuguese-language consultation service (https://ciberduvidas.iscte-iul.pt). For each of the 1000 original questions, five paraphrased reformulations were generated using the mistral-large-2512 model via the Mistral API, according to five distinct query profiles designed to simulate the range of phrasings a real user might submit: Synthetic: short, direct, search-engine style queries- Formal: grammatically rigorous and highly detailed queries- Informal: casual, conversational tone queries- Professor: queries using technical pedagogical/linguistic terminology- Student: direct, comprehension-focused queries This scaled benchmark served as a primary evaluation suite for the retrieval and reranking components of a Conversational Agent system. It was used to: compare embedding models, retrieval strategies, evaluate a domain-adapted cross-encoder reranker and support an automatic Ragas-based comparison of candidate generation models. Each entry includes:- id: unique identifier of the original Ciberdúvidas question- url: direct link to the original question on the Ciberdúvidas website- paraphrases: object with the five paraphrased query variants (synthetic, formal, informal, teacher, student) | eng |
| dc.identifier.citation | Moura, P., Batista, F., & Lopes, A. (2026). Ciberdúvidas da Língua Portuguesa: European Portuguese Linguistics Paraphrase Benchmark - 1000-Question Scaled Evaluation Set (5000 Queries) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.22819697 | |
| dc.identifier.doi | https://doi.org/10.5281/zenodo.22819697 | |
| dc.identifier.uri | https://hdl.handle.net/10071/38671 | |
| dc.language.iso | eng | |
| dc.rights | open access | |
| dc.rights.license | cc-by-4.0 | |
| dc.subject | European Portuguese | eng |
| dc.subject | Information Storage and Retrieval | eng |
| dc.subject | Natural Lenguage Processing | eng |
| dc.title | Ciberdúvidas da Língua Portuguesa: European Portuguese Linguistics Paraphrase Benchmark - 1000-Question Scaled Evaluation Set (5000 Queries) | eng |
| dc.type | dataset | |
| dspace.entity.type | Publication |
Ficheiros
Pacote original
1 - 1 de 1
A carregar...
- Nome:
- dataset_hdl_38671.pdf
- Tamanho:
- 248.09 KB
- Formato:
- Adobe Portable Document Format
