A small experiment using both Mecab and Tinysegmenter to create a tokenized list of Japanese sentences in JSON, taken from the Tatoeba corpus.
-
Updated
Mar 25, 2021 - Python
A small experiment using both Mecab and Tinysegmenter to create a tokenized list of Japanese sentences in JSON, taken from the Tatoeba corpus.
An app for automating the creation of cloze (fill-in-the-blank) vocabulary and grammar activities. Powered by the Tatoeba corpus.
Katakana loanword (外来語) analysis on Tatoeba v2022-03-03 covering data cleaning, visualization, and frequency analysis in Python.
To associate your repository with the tatoeba-corpus topic, visit your repo's landing page and select "manage topics."