OLAC Record
oai:www.clarin.si:11356/1059

Metadata
Title:Serbian-English parallel corpus srenWaC 1.0
Bibliographic Citation:http://hdl.handle.net/11356/1059
Creator:Ljubešić, Nikola
Esplà-Gomis, Miquel
Ortiz Rojas, Sergio
Klubička, Filip
Toral, Antonio
Date (W3CDTF):2016-03-09T16:51:44Z
Date Available:2016-03-09T16:51:44Z
Description:The srenWaC corpus consists of sentence aligned parallel Serbian-English texts crawled from the .rs top-level domain for Serbia. The corpus was built with Spidextor (https://github.com/abumatran/spidextor), a tool that glues together the output of SpiderLing used for crawling and Bitextor used for bitext extraction. The accuracy of the extracted bitext, given the evaluation results on other languages, can be estimated at 74% on the sentence level and 76% on the word level.
Identifier (URI):http://hdl.handle.net/11356/1059
Language:Serbian
English
Language (ISO639):srp
eng
Publisher:Jožef Stefan Institute
Rights:CLARIN.SI User Licence for Internet Corpora
http://www.clarin.si/info/wp-content/uploads/2016/01/CLARIN.SI-WAC-2016-01.pdf
Subject:parallel corpus
web corpus
multilingual
Type:corpus
Type (DCMI):Text
Type (OLAC):primary_text

OLAC Info

Archive:  Slovenian language resource repository CLARIN.SI
Description:  http://www.language-archives.org/archive/clarin.si
GetRecord:  OAI-PMH request for OLAC format
GetRecord:  Pre-generated XML file

OAI Info

OaiIdentifier:  oai:www.clarin.si:11356/1059
DateStamp:  2018-10-29
GetRecord:  OAI-PMH request for simple DC format

Search Info

Citation: Ljubešić, Nikola; Esplà-Gomis, Miquel; Ortiz Rojas, Sergio; Klubička, Filip; Toral, Antonio. 2016. Jožef Stefan Institute.
Terms: area_Europe country_GB country_RS dcmi_Text iso639_eng iso639_srp olac_primary_text


http://www.language-archives.org/item.php/oai:www.clarin.si:11356/1059
Up-to-date as of: Mon Jun 10 9:21:33 EDT 2019