OLAC Record
oai:www.ldc.upenn.edu:LDC2016T18

Metadata
Title:ARL Arabic Dependency Treebank
Access Rights:Licensing Instructions for Subscription & Standard Members, and Non-Members: http://www.ldc.upenn.edu/language-resources/data/obtaining
Bibliographic Citation:Tratz, Stephen. ARL Arabic Dependency Treebank LDC2016T18. Web Download. Philadelphia: Linguistic Data Consortium, 2016
Contributor:Tratz, Stephen C
Date (W3CDTF):2016
Date Issued (W3CDTF):2016-09-15
Description:*Introduction* ARL Arabic Dependency Treebank was developed by the US Army Research Laboratory (ARL) and was derived from four LDC resources: Arabic Treebank (ATB) Part 1 v 4.1 (LDC2010T13), Part 2 v 3.1 (LDC2011T09), Part 3 v 3.2 (LDC2010T08) and Broadcast News v 1.0 (LDC2012T07). LDC's ATB series follows the constituency or phrase structure approach to treebank development in which clauses are divided into noun phrases and verb phrases and in each sentence, one or more nodes may correspond to one element. Dependency grammar, on the other hand, is based on the idea that the verb is the center of the clause structure and that other units in the sentence are connected to the verb as directed links or dependencies. This is a one-to-one correspondence: for every element in the sentence there is one node in the sentence structure that corresponds to that element. ARL Arabic Dependency Treebank was generated using constituency-to-dependency software written at ARL. *Data* The source data in this release consists of Arabic newswire and broadcast programming collected by LDC from various news and broadcast providers. The files are in an 11-column tab-separated format with one or more blank lines between sentences. All files are UTF-8 encoded. Further information about the corpus structure is contained in the documentation accompanying this release. *Samples* Please view this sample. *Updates* None at this time.
Extent:Corpus size: 258272 KB
Identifier:LDC2016T18
https://catalog.ldc.upenn.edu/LDC2016T18
ISBN: 1-58563-769-6
ISLRN: 110-972-255-734-6
Language:Standard Arabic
English
Language (ISO639):arb
eng
License:LDC User Agreement for Non-Members: https://catalog.ldc.upenn.edu/license/ldc-non-members-agreement.pdf
Medium:Distribution: Web Download
Publisher:Linguistic Data Consortium
Publisher (URI):https://www.ldc.upenn.edu
Relation (URI):https://catalog.ldc.upenn.edu/docs/LDC2016T18
Rights Holder:Portions © 2007-2008 Abu Dhabi TV, © 2000 Agence France Presse, © 2007 Al Alam News Channel, © 2007-2008 Al Arabiya, © 2008 Al Baghdadya TV, © 2008 Al Fayha, © 2007-2008 Al Iraqiyah, © 2005-2007 Aljazeera, © 2007 Al Ordiniyah, © 2008 Al Sharqiyah, © 2002 An Nahar, © 2005, 2007-2008 Dubai TV, © 2007 Kuwait TV, © 2008 Oman TV, © 2007 PAC Ltd, © 2007-2008 Saudi TV, © 2007 Syria TV, © 2001-2002 Ummah Press, © 2003, 2004, 2005, 2007, 2008, 2009, 2010, 2011, 2012, 2016 Trustees of the University of Pennsylvania
Type (DCMI):Text
Type (OLAC):primary_text

OLAC Info

Archive:  The LDC Corpus Catalog
Description:  http://www.language-archives.org/archive/www.ldc.upenn.edu
GetRecord:  OAI-PMH request for OLAC format
GetRecord:  Pre-generated XML file

OAI Info

OaiIdentifier:  oai:www.ldc.upenn.edu:LDC2016T18
DateStamp:  2019-01-03
GetRecord:  OAI-PMH request for simple DC format

Search Info

Citation: Tratz, Stephen C. 2016. Linguistic Data Consortium.
Terms: area_Asia area_Europe country_GB country_SA dcmi_Text iso639_arb iso639_eng olac_primary_text


http://www.language-archives.org/item.php/oai:www.ldc.upenn.edu:LDC2016T18
Up-to-date as of: Sat Jul 27 9:56:54 EDT 2019