OLAC Record: LT Corpus

OLAC Record
oai:catalogue.elra.info:ELRA-W0059

Metadata

Title: LT Corpus

Access Rights: Rights available for: nonCommercialUse, commercialUse

Date Available (W3CDTF): 2012-12-05

Date Issued (W3CDTF): 2012-12-05

Date Modified (W3CDTF): 2012-12-05

Description: The LT Corpus is composed of 70 fiction texts from Portuguese renowned authors. The corpus contains 1,781,083 tokens. The texts date from before 1940.The corpus is delivered in one file, in two different formats. The txt version has one sentence per line, an identification number for each text and no further annotation. The cqpweb file is one token per line, followed by pos tag and lemma, and is annotated for NP chunks. The LT Corpus is a copyright free subset of the Corpus of Reference of Contemporary Portuguese and follows the same annotation scheme. For more information on its preparation and annotation, see: Généreux, M., I. Hendrickx, A. Mendes (2012) “A Large Portuguese Corpus On-Line: Cleaning and Preprocessing”. In Caseli, H. et al. (eds.) Computational Processing of the Portuguese Language. Proceedings of the 10th International Conference PROPOR1012. Berlin, Heidelberg: Springer-Verlag, pp. 113-120. The Corpus is delivered with the annotation manual of the CRPC, a metadata file and a narrative description of the resource.

Identifier: ELRA-W0059

ISLRN: 569-208-468-863-2

Identifier (URI): https://catalog.elra.info/en-us/repository/browse/ELRA-W0059/

Language: Portuguese

Language (ISO639): por

Medium: Not specified

Publisher: ELRA (European Language Resources Association)

Type (DCMI): Text

Type (OLAC): primary_text

OLAC Info

Archive: ELRA Catalogue of Language Resources

Description: http://www.language-archives.org/archive/catalogue.elra.info

GetRecord: OAI-PMH request for OLAC format

GetRecord: Pre-generated XML file

OAI Info

OaiIdentifier: oai:catalogue.elra.info:ELRA-W0059

DateStamp: 2012-12-05

GetRecord: OAI-PMH request for simple DC format

Search Info
Citation: n.a. 2012. ELRA (European Language Resources Association).
Terms: area_Europe country_PT dcmi_Text iso639_por olac_primary_text

http://www.language-archives.org/item.php/oai:catalogue.elra.info:ELRA-W0059
Up-to-date as of: Wed Jul 15 7:04:38 EDT 2026

Metadata
Title:		LT Corpus
Access Rights:		Rights available for: nonCommercialUse, commercialUse
Date Available (W3CDTF):		2012-12-05
Date Issued (W3CDTF):		2012-12-05
Date Modified (W3CDTF):		2012-12-05
Description:		The LT Corpus is composed of 70 fiction texts from Portuguese renowned authors. The corpus contains 1,781,083 tokens. The texts date from before 1940.The corpus is delivered in one file, in two different formats. The txt version has one sentence per line, an identification number for each text and no further annotation. The cqpweb file is one token per line, followed by pos tag and lemma, and is annotated for NP chunks. The LT Corpus is a copyright free subset of the Corpus of Reference of Contemporary Portuguese and follows the same annotation scheme. For more information on its preparation and annotation, see: Généreux, M., I. Hendrickx, A. Mendes (2012) “A Large Portuguese Corpus On-Line: Cleaning and Preprocessing”. In Caseli, H. et al. (eds.) Computational Processing of the Portuguese Language. Proceedings of the 10th International Conference PROPOR1012. Berlin, Heidelberg: Springer-Verlag, pp. 113-120. The Corpus is delivered with the annotation manual of the CRPC, a metadata file and a narrative description of the resource.
Identifier:		ELRA-W0059
Identifier:		ISLRN: 569-208-468-863-2
Identifier (URI):		https://catalog.elra.info/en-us/repository/browse/ELRA-W0059/
Language:		Portuguese
Language (ISO639):		por
Medium:		Not specified
Publisher:		ELRA (European Language Resources Association)
Type (DCMI):		Text
Type (OLAC):		primary_text
OLAC Info
Archive:		ELRA Catalogue of Language Resources
Description:		http://www.language-archives.org/archive/catalogue.elra.info
GetRecord:		OAI-PMH request for OLAC format
GetRecord:		Pre-generated XML file
OAI Info
OaiIdentifier:		oai:catalogue.elra.info:ELRA-W0059
DateStamp:		2012-12-05
GetRecord:		OAI-PMH request for simple DC format
Search Info
Citation:		n.a. 2012. ELRA (European Language Resources Association).
Terms:		area_Europe country_PT dcmi_Text iso639_por olac_primary_text