GitHub - CUNY-CL/wikipron: Massively multilingual pronunciation mining (original) (raw)
WikiPron is a command-line tool and Python API for mining multilingual pronunciation data from Wiktionary, as well as a database of pronunciation dictionaries mined using this tool.
If you use WikiPron in your research, please cite the following:
Jackson L. Lee, Lucas F.E. Ashby, M. Elizabeth Garza, Yeonju Lee-Sikka, Sean Miller, Alan Wong, Arya D. McCarthy, and Kyle Gorman (2020). Massively multilingual pronunciation mining with WikiPron. In Proceedings of the 12th Language Resources and Evaluation Conference, pages 4223-4228. [bibtex]
Command-line tool
Installation
Usage
Quick start
After installation, the terminal command wikipron will be available. As a basic example, the following command scrapes G2P data for French:
Specifying the language
The language is indicated by a three-letter ISO 639-3 language code, e.g., fra for French. For which languages can be scraped,hereis the complete list of languages on Wiktionary that have pronunciation entries.
Specifying the dialect
One can optionally specify dialects to target using the --dialect flag. The dialect name can be found together with the transcription on Wiktionary. For example, "(UK, US) IPA: /təˈmɑːtəʊ/". To restrict to the union of dialects use the pipe character '|': e.g., --dialect='General American | US'. Transcriptions which lack a dialect specification are selected regardless of the value of this flag.
Specifying the transcription level
By default, WikiPron selects broad pronunciations in angled brackets /like this/. One can instead select narrow transcriptions written [like this] using the --narrow flag. Note that some languages only have broad or narrow transcriptions (e.g., Russian only has the latter.
Segmentation
By default, the segments library is used to segment the transcription into whitespace. The segmentation tends to place IPA diacritics and modifiers on the "parent" symbol. For instance, [kʰæt] is rendered kʰ æ t. This can be disabled using the --no-segment flag.
Parentheses
Some transcriptions contain parentheses to indicate optional sounds (e.g., English A&E /eɪ.ən(d)ˈiː/). The --parens flag controls how they are handled:expand (default) generates all variants, skip removes parentheses and their content, and show keeps parentheses as-is in the output.
Output
The scraped data is organized with each <word, pronunciation> pair on its own line, where the word and pronunciation are separated by a tab. Note that the pronunciation is in International Phonetic Alphabet (IPA), segmented by spaces that correctly handle the combining and modifier diacritics for modeling purposes, e.g., we have kʰ æ t with the aspirated k instead ofk ʰ æ t.
For illustration, here is a snippet of French data scraped by WikiPron:
accrémentitielle a k ʁ e m ɑ̃ t i t j ɛ l accrescent a k ʁ ɛ s ɑ̃ accrétion a k ʁ e s j ɔ̃ accrétions a k ʁ e s j ɔ̃
By default, the scraped data appears in the terminal. To save the data in a TSV file, please redirect the standard output to a filename of your choice:
Advanced options
The wikipron terminal command has an array of options to configure your scraping run. For a full list of the options, please run wikipron -h.
Python API
The underlying module can also be used from Python. A standard workflow looks like:
import wikipron
config = wikipron.Config(key="fra") # French, with default options. for word, pron in wikipron.scrape(config): ...
Data
We also make available a database of over 3 million word/pronunciation pairs mined using WikiPron.
Models
We host grapheme-to-phoneme models and modeling software in a separate repository.
Development
Repository
The source code of WikiPron is hosted on GitHub athttps://github.com/CUNY-CL/wikipron, where development also happens.
For the latest changes not yet released through pip or working on the codebase yourself, you may obtain the latest source code through GitHub and git:
- Create a fork of the
wikipronrepo on your GitHub account. - Clone from your fork:
git clone https://github.com//wikipron.git
cd wikipron - Set up a Python virtual environment. We recommend using uv:
uv python install 3.14
uv venv --python 3.14
source .venv/bin/activate - Install WikiPron in the "editable" mode together with the core and dev dependencies:
uv pip install -e ".[dev]"
We keep track of notable changes inCHANGELOG.md.
Contributing
For questions, bug reports, and feature requests, please file an issue.
If you would like to contribute to the wikipron codebase, please seeCONTRIBUTING.md.
License
WikiPron is released under an Apache 2.0 license. Please seeLICENSE.txt for details.
Please note that Wiktionary data in thedata/ directory hasits own licensing terms.