Corpus Linguistics for Korean with Universal Dependencies: Concepts and Hands-On Practice
Date & Time: Thursday, August 6, 9:30am – 12:00pm
Learning objectives:
- Participants will learn to:
- Understand the core principles of corpus linguistics and Universal Dependencies (UD)
- Navigate and query UD corpora (e.g., extract frequencies)
- Compare morphosyntactic patterns in L1 (UD_Korean-GSD) vs. L2 (UD_Korean-KSL) corpora
- UD_Korean-GSD — L1 Korean corpus (sourced from online news and blog posts): https://github.com/UniversalDependencies/UD_Korean-GSD
- UD_Korean-KSL — L2 Korean corpus (sourced from argumentative and narrative essays): https://github.com/UniversalDependencies/UD_Korean-KSL
- Report UD-based findings in academic contexts
Survey: If you plan to attend the Korean workshop, please complete this short survey: https://forms.gle/S4C7sXGvWB8Uvhjy9
Click here to view the tentative schedule
Schedule (subject to change) ***Please bring your own devices***
- Opening
- Workshop goals and structure
- Quick gauge of familiarity with Korean morphosyntax and corpus tools
- Brief introduction to Corpus Linguistics
- Universal Dependencies
- Overview of UD’s goals and design
- Elements of UD annotation
- Why UD works for workshop
- Available Korean UD corpora
- Group activity: Annotation decisions and criteria
- Hands-on activities
- UD annotation and frequency analysis: Annotate texts using available NLP tools and extract linguistic feature frequencies
- Concordance and dependency patterns: Identify and interpret morphosyntactic patterns using UD-based concordance data
- Examine recurring patterns in sample texts
- Closing
- How to report and disseminate corpus-based findings (e.g., avoiding overclaims)
- Introduce selected conferences, journals, and other venues for presenting results
An Introduction to Japanese Spoken Corpora and Tools
Date & Time: Thursday, August 6, 3:30-6:00 pm
Description: This workshop introduces Japanese spoken corpora available through the National Institute for Japanese Language and Linguistics (NINJAL), as well as corpus search tools such as Kotonoha and Chunagon. It presents sample studies using one or more of these corpora and provides hands-on practice in searching corpus data and analyzing the results. The workshop is intended for students and researchers who work with or are interested in working with Japanese spoken data and who have little or no experience with NINJAL corpora and corpus search tools. The workshop is not intended for advanced users.
Survey: If you plan to attend the workshop, please complete the short survey at the link below by July 24. The workshop is open to and free for JK33 conference registrants.
Japanese Pre-conference Workshop Registration and Survey Form