OLAC Record: The AQUAINT Corpus of English News Text

OLAC Record
oai:www.ldc.upenn.edu:LDC2002T31

Metadata

Title: The AQUAINT Corpus of English News Text

Access Rights: Licensing Instructions for Subscription & Standard Members, and Non-Members: http://www.ldc.upenn.edu/language-resources/data/obtaining

Bibliographic Citation: Graff, David. The AQUAINT Corpus of English News Text LDC2002T31. Web Download. Philadelphia: Linguistic Data Consortium, 2002

Contributor: Graff, David

Date (W3CDTF): 2002

Date Issued (W3CDTF): 2002-09-26

Description: *Introduction* The AQUAINT Corpus, Linguistic Data Consortium (LDC) catalog number LDC2002T31 and ISBN 1-58563-240-6 consists of newswire text data in English, drawn from three sources: the Xinhua News Service (People's Republic of China), the New York Times News Service, and the Associated Press Worldstream News Service. It was prepared by the LDC for the AQUAINT Project, and will be used in official benchmark evaluations conducted by National Institute of Standards and Technology (NIST). *Data* The data files contain roughly 375 million words correlating to about 3GB of data. The text data are separated into directories by source (apw, nyt, xie); within each source, data files are subdivided by year, and within each year, there is one file per date of collection. Each file is named to reflect the source and date, and contains a stream of SGML-tagged text data presenting the series of news stories reported on the given date as a concatenation of DOC elements (i.e. blocks of text bounded by and tags). All data files are published in compressed form, using the GNU "gzip" utility; as such, all files have a ".gz" extension, and will have null file name extension when uncompressed in the usual way (i.e. just the base file name, consisting of "YYYYMMDD_SRC"). While all the data files are covered by a single DTD, it is not the case that they all have a single pattern of markup. Rather, all files share a core markup structure, with minor variations in the peripheral regions of each DOC element, and the DTD has been written to accommodate the variations. *Updates* 19980614_NYT.gz was left off in the conversion from CD to DVD. An update was issued on 09/13/2012. All copies ordered after this date will be complete. Contact ldc@ldc.upenn.edu for more information.

Extent: Corpus size: 2202009 KB

Identifier: LDC2002T31

https://catalog.ldc.upenn.edu/LDC2002T31

ISBN: 1-58563-240-6

ISLRN: 153-002-267-999-9

DOI: 10.35111/pcbv-jq63

Language: English

Language (ISO639): eng

License: LDC User Agreement for Non-Members: https://catalog.ldc.upenn.edu/license/ldc-non-members-agreement.pdf

Medium: Distribution: Web Download

Publisher: Linguistic Data Consortium

Publisher (URI): https://www.ldc.upenn.edu

Relation (URI): https://catalog.ldc.upenn.edu/docs/LDC2002T31

Rights Holder: Portions © 1998-2000 New York Times, © 1998-2000 The Associated Press, © 1996-2000 Xinhua News Agency, © 2002 Trustees of the University of Pennsylvania

Type (DCMI): Text

Type (OLAC): primary_text

OLAC Info

Archive: The LDC Corpus Catalog

Description: http://www.language-archives.org/archive/www.ldc.upenn.edu

GetRecord: OAI-PMH request for OLAC format

GetRecord: Pre-generated XML file

OAI Info

OaiIdentifier: oai:www.ldc.upenn.edu:LDC2002T31

DateStamp: 2020-11-30

GetRecord: OAI-PMH request for simple DC format

Search Info
Citation: Graff, David. 2002. Linguistic Data Consortium.
Terms: area_Europe country_GB dcmi_Text iso639_eng olac_primary_text

http://www.language-archives.org/item.php/oai:www.ldc.upenn.edu:LDC2002T31
Up-to-date as of: Wed Oct 29 7:00:13 EDT 2025

Metadata
Title:		The AQUAINT Corpus of English News Text
Access Rights:		Licensing Instructions for Subscription & Standard Members, and Non-Members: http://www.ldc.upenn.edu/language-resources/data/obtaining
Bibliographic Citation:		Graff, David. The AQUAINT Corpus of English News Text LDC2002T31. Web Download. Philadelphia: Linguistic Data Consortium, 2002
Contributor:		Graff, David
Date (W3CDTF):		2002
Date Issued (W3CDTF):		2002-09-26
Description:		Introduction The AQUAINT Corpus, Linguistic Data Consortium (LDC) catalog number LDC2002T31 and ISBN 1-58563-240-6 consists of newswire text data in English, drawn from three sources: the Xinhua News Service (People's Republic of China), the New York Times News Service, and the Associated Press Worldstream News Service. It was prepared by the LDC for the AQUAINT Project, and will be used in official benchmark evaluations conducted by National Institute of Standards and Technology (NIST). Data The data files contain roughly 375 million words correlating to about 3GB of data. The text data are separated into directories by source (apw, nyt, xie); within each source, data files are subdivided by year, and within each year, there is one file per date of collection. Each file is named to reflect the source and date, and contains a stream of SGML-tagged text data presenting the series of news stories reported on the given date as a concatenation of DOC elements (i.e. blocks of text bounded by and tags). All data files are published in compressed form, using the GNU "gzip" utility; as such, all files have a ".gz" extension, and will have null file name extension when uncompressed in the usual way (i.e. just the base file name, consisting of "YYYYMMDD_SRC"). While all the data files are covered by a single DTD, it is not the case that they all have a single pattern of markup. Rather, all files share a core markup structure, with minor variations in the peripheral regions of each DOC element, and the DTD has been written to accommodate the variations. Updates 19980614_NYT.gz was left off in the conversion from CD to DVD. An update was issued on 09/13/2012. All copies ordered after this date will be complete. Contact ldc@ldc.upenn.edu for more information.
Extent:		Corpus size: 2202009 KB
Identifier:		LDC2002T31
		https://catalog.ldc.upenn.edu/LDC2002T31
		ISBN: 1-58563-240-6
		ISLRN: 153-002-267-999-9
		DOI: 10.35111/pcbv-jq63
Language:		English
Language (ISO639):		eng
License:		LDC User Agreement for Non-Members: https://catalog.ldc.upenn.edu/license/ldc-non-members-agreement.pdf
Medium:		Distribution: Web Download
Publisher:		Linguistic Data Consortium
Publisher (URI):		https://www.ldc.upenn.edu
Relation (URI):		https://catalog.ldc.upenn.edu/docs/LDC2002T31
Rights Holder:		Portions © 1998-2000 New York Times, © 1998-2000 The Associated Press, © 1996-2000 Xinhua News Agency, © 2002 Trustees of the University of Pennsylvania
Type (DCMI):		Text
Type (OLAC):		primary_text
OLAC Info
Archive:		The LDC Corpus Catalog
Description:		http://www.language-archives.org/archive/www.ldc.upenn.edu
GetRecord:		OAI-PMH request for OLAC format
GetRecord:		Pre-generated XML file
OAI Info
OaiIdentifier:		oai:www.ldc.upenn.edu:LDC2002T31
DateStamp:		2020-11-30
GetRecord:		OAI-PMH request for simple DC format
Search Info
Citation:		Graff, David. 2002. Linguistic Data Consortium.
Terms:		area_Europe country_GB dcmi_Text iso639_eng olac_primary_text