OLAC Record
oai:catalogue.elra.info:ELRA-W0060

Metadata
Title:PTPARL Corpus
Abstract:The PTPARL Corpus contains 1,076 texts consisting of adapted transcriptions of the Portuguese Parliament sessions. The corpus contains 1,000,441 tokens. The corpus is delivered in one file, in two different formats. The txt version has one sentence per line, an identification number for each text and no further annotation. The cqpweb file is one token per line, followed by pos tag and lemma, and is annotated for NP chunks.
Access Rights:Rights available for: Commercial Use, Research Use
Date Available (W3CDTF):2012-12-05
Date Issued (W3CDTF):2012-12-05
Date Modified (W3CDTF):2012-12-05
Description:Written Corpora
The PTPARL Corpus contains 1,076 texts consisting of adapted transcriptions of the Portuguese Parliament sessions. The corpus contains 1,000,441 tokens. The corpus is delivered in one file, in two different formats. The txt version has one sentence per line, an identification number for each text and no further annotation. The cqpweb file is one token per line, followed by pos tag and lemma, and is annotated for NP chunks. The PTPARL Corpus is a subset of the Corpus of Reference of Contemporary Portuguese and follows the same annotation scheme. For more information on its preparation and annotation, see: G?n?reux, M., I. Hendrickx, A. Mendes (2012) ?A Large Portuguese Corpus On-Line: Cleaning and Preprocessing?. In Caseli, H. et al. (eds.) Computational Processing of the Portuguese Language. Proceedings of the 10th International Conference PROPOR1012. Berlin, Heidelberg: Springer-Verlag, pp. 113-120. The Corpus is delivered with the annotation manual of the CRPC corpus, a metadata file and a narrative description of the resource.
Identifier:ELRA-W0060
http://catalog.elra.info/product_info.php?products_id=1179
Language:Portuguese
Language (ISO639):por
Publisher:ELRA (European Language Resources Association)
Type (DCMI):Text
Type (OLAC):primary_text

OLAC Info

Archive:  ELRA Catalogue of Language Resources
Description:  http://www.language-archives.org/archive/catalogue.elra.info
GetRecord:  OAI-PMH request for OLAC format
GetRecord:  Pre-generated XML file

OAI Info

OaiIdentifier:  oai:catalogue.elra.info:ELRA-W0060
DateStamp:  2012-12-05
GetRecord:  OAI-PMH request for simple DC format

Search Info

Citation: n.a. 2012. ELRA (European Language Resources Association).
Terms: area_Europe country_PT dcmi_Text iso639_por olac_primary_text


http://www.language-archives.org/item.php/oai:catalogue.elra.info:ELRA-W0060
Up-to-date as of: Fri Jun 23 1:06:23 EDT 2017