Authors:

Niladri Sekhar Dash ⁰,
S. Arulmozi ¹

Niladri Sekhar Dash
1. Linguistic Research Unit, Indian Statistical Institute, Kolkata, India
View author publications

You can also search for this author in PubMed Google Scholar
S. Arulmozi
1. Centre for Applied Linguistics and Translation Studies, University of Hyderabad, Hyderabad, India
View author publications

You can also search for this author in PubMed Google Scholar

Presents discussions in simple English to cater to the needs of non-native English readers
Is a must-read for graduate and postgraduate courses on applied and computational linguistics
Includes many pedagogical tools for easy comprehension of technical discussions

6122 Accesses
7 Citations
1 Altmetric

Buy it now

eBook USD 69.99

Price excludes VAT (USA)

Softcover Book USD 89.99

Price excludes VAT (USA)

Hardcover Book USD 89.99

Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Other ways to access

Licence this eBook for your library

Learn about institutional subscriptions

This is a preview of subscription content, log in via an institution to check for access.

Table of contents (15 chapters)

Front Matter

Pages i-xxix

PDF
Definition of ‘Corpus’
- Niladri Sekhar Dash, S. Arulmozi
Pages 1-15
Features of a Corpus
- Niladri Sekhar Dash, S. Arulmozi
Pages 17-34
Genre of Text
- Niladri Sekhar Dash, S. Arulmozi
Pages 35-49
Nature of Data
- Niladri Sekhar Dash, S. Arulmozi
Pages 51-65
Type and Purpose of Text
- Niladri Sekhar Dash, S. Arulmozi
Pages 67-83
Nature of Text Application
- Niladri Sekhar Dash, S. Arulmozi
Pages 85-99
Parallel Translation Corpus
- Niladri Sekhar Dash, S. Arulmozi
Pages 101-124
Web Text Corpus
- Niladri Sekhar Dash, S. Arulmozi
Pages 125-146
Pre-digital Corpora (Part 1)
- Niladri Sekhar Dash, S. Arulmozi
Pages 147-165
Pre-digital Corpora (Part 2)
- Niladri Sekhar Dash, S. Arulmozi
Pages 167-186
Digital Text Corpora (Part 1)
- Niladri Sekhar Dash, S. Arulmozi
Pages 187-202
Digital Text Corpora (Part 2)
- Niladri Sekhar Dash, S. Arulmozi
Pages 203-219
Digital Speech Corpora
- Niladri Sekhar Dash, S. Arulmozi
Pages 221-239
Utilization of Language Corpora
- Niladri Sekhar Dash, S. Arulmozi
Pages 241-258
Limitations of Language Corpora
- Niladri Sekhar Dash, S. Arulmozi
Pages 259-272
Back Matter

Pages 273-293

PDF

About this book

This book discusses key issues of corpus linguistics like the definition of the corpus, primary features of a corpus, and utilization and limitations of corpora. It presents a unique classification scheme of language corpora to show how they can be studied from the perspective of genre, nature, text type, purpose, and application. A reference to parallel translation corpus is mandatory in the discussion of corpus generation, which the authors thoroughly address here, with a focus on Indian language corpora and English. Web-text corpus, a new development in corpus linguistics, is also discussed with elaborate reference to Indian web text corpora. The book also presents a short history of corpus generation and provides scenarios before and after the advent of computer-generated digital corpora.

This book has several important features: it discusses many technical issues of the field in a lucid manner; contains extensive new diagrams and chartsfor easy comprehension; and presents discussions in simplified English to cater to the needs of non-native English readers. This is an important resource authored by academics who have many years of experience teaching and researching corpus linguistics. Its focus on Indian languages and on English corpora makes it applicable to students of graduate and postgraduate courses in applied linguistics, computational linguistics and language processing in South Asia and across countries where English is spoken as a first or second language.

Keywords

Authors and Affiliations

Linguistic Research Unit, Indian Statistical Institute, Kolkata, India

Niladri Sekhar Dash
Centre for Applied Linguistics and Translation Studies, University of Hyderabad, Hyderabad, India

S. Arulmozi

About the authors

Niladri Sekhar Dash, PhD, is Associate Professor in the Linguistic Research Unit of the Indian Statistical Institute, Kolkata. He has been working on Corpus Linguistics, Language Technology, Natural Language Processing, Language Documentation and Digitization, Computational Lexicography, Computer Assisted Language Teaching, and Manual and Machine Translation for over two decades. He has to his credit 15 research monographs and 160 research papers in peer-reviewed international and national journals, anthologies, and conference proceedings. He has delivered lectures and taught courses as an invited scholar at more than 30 universities and institutes in India and abroad. He has acted as a consultant for several organizations working on Language Technology and Natural Language Processing. Dr. Dash is the Principal Investigator for 5 language technology projects funded by the Government of India and the Indian Statistical Institute, Kolkata. He is the Editor-in-Chief of the Journal of Advanced Linguistic Studies – a peer-reviewed international journal of linguistics; and Editorial Board Member of 5 international journals. He is member of several linguistics associations across the world and a regular Ph.D. thesis adjudicator for several Indian universities. At present Dr. Dash is working on a Digital Pronunciation Dictionary for Bangla, Hindi-Bangla Parallel Translation Corpus Generation, Endangered Language Documentation and Digitization, POS Tagging and Chunking, Word Sense Disambiguation, Manual and Machine Translation, and Computer Assisted Language Teaching, etc. Details of Dr. Dash are at https://sites.google.com/site/nsdashisi/home/

S. Arulmozi, PhD, is Assistant Professor at the Centre for Applied Linguistics and Translation Studies (CALTS), University of Hyderabad, India. He has previously taught at the Dravidian University, Kuppam; acted as Guest Faculty at CALTS, University of Hyderabad; worked as Research Staff at the Anna University, Chennai; as Project Fellow at the Tamil University, Thanjavur; and as Language Assistant-Tamil at the Central Institute of Indian Languages, Mysore. Dr. Arulmozi has been working on Corpus Linguistics for some years and has been trained professionally in WordNet. He has successfully carried out projects on Corpus Linguistics and WordNet funded by the Government of India and has also conducted a workshop on language technology at the University of Malaya, Kuala Lumpur, Malaysia. To his credit, he has 1 research monograph and 15 research papers in peer-reviewed international and national journals, edited volumes, and conference proceedings.

Bibliographic Information

Book Title: History, Features, and Typology of Language Corpora
Authors: Niladri Sekhar Dash, S. Arulmozi
DOI: https://doi.org/10.1007/978-981-10-7458-5
Publisher: Springer Singapore
eBook Packages: Social Sciences, Social Sciences (R0)
Copyright Information: Springer Nature Singapore Pte Ltd. 2018
Hardcover ISBN: 978-981-10-7457-8Published: 12 February 2018
Softcover ISBN: 978-981-13-5638-4Published: 11 February 2019
eBook ISBN: 978-981-10-7458-5Published: 01 February 2018
Edition Number: 1
Number of Pages: XXIX, 293
Number of Illustrations: 58 b/w illustrations, 17 illustrations in colour
Topics: Corpus Linguistics, Natural Language Processing (NLP), Language Teaching

Publish with us

Policies and ethics

Authors:

Sections

Buy it now

Buying options

Other ways to access

Table of contents (15 chapters)

Front Matter

Back Matter

About this book

Keywords

Authors and Affiliations

Linguistic Research Unit, Indian Statistical Institute, Kolkata, India

Centre for Applied Linguistics and Translation Studies, University of Hyderabad, Hyderabad, India

About the authors

Bibliographic Information

Publish with us

Buy it now

Buying options

Other ways to access

Search

Navigation