Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Improving a Dictionary-Based Minimum Forward Matching Tokenizer to Better Handle Word Segmentation of Japanese Sentences
Blekinge Institute of Technology, Faculty of Computing, Department of Computer Science.
2025 (English)Independent thesis Basic level (university diploma), 10 credits / 15 HE creditsStudent thesis
Abstract [en]

This thesis addresses the challenge of accurately tokenizing Japanese sentences by improving a sole trader’s Python-based sentence-splitter system to better handle the splitting of Japanese sentences. Japanese poses unique difficulties during tokenization due to its agglutinative nature and due to the fact that spaces are not used to delimit words. We evaluate the baseline system, having a token lexicon of some 120 000 entries, against a Short Unit Word gold standard consisting of a sample of 10 000 sentences from Tatoeba, splitted using fugashi and the UniDic token lexicon. Precision, recall and F1-scores are calculated and an F1-score of 0.6537 is established for the baseline system. We enhance the token lexicon by (1) adding root verb and root adjective stems, (2) adding five additional verb stems that are used in UniDic and (3) incorporating 5 050 high-frequency tokens missing from the original lexicon. These additions raise the F1-score to 0.7235, demonstrating that while token lexicon expansion can yield some improvements in F1-score it is ultimately insufficient to reach the industry standard (96-99 %) without either pruning non-UniDic entries from the customer’s pre-existing token lexicon or by altering the phrase matching algorithm in the customer’s sentence-splitter system (which is currently a Minimum Forward Matching method). It is concluded that, while a base of Short Unit Words in the customer’s token lexicon is ultimately necessary, a move to a Long Unit Word standard may allow the customer to reach an industry standard, or somewhere close to it, while preserving the pre-existing data in the customer’s token lexicon.

Place, publisher, year, edition, pages
2025. , p. 29
Keywords [en]
tokenization, japanese, word segmentation, tokenizer
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:bth-28021OAI: oai:DiVA.org:bth-28021DiVA, id: diva2:1981852
External cooperation
The Swedish Polyglot
Subject / course
PA1438 Självständigt arbete Webbprogrammering
Educational program
PAGWG Webbprogrammering
Supervisors
Examiners
Available from: 2025-08-05 Created: 2025-07-06 Last updated: 2025-09-30Bibliographically approved

Open Access in DiVA

fulltext(336 kB)84 downloads
File information
File name FULLTEXT01.pdfFile size 336 kBChecksum SHA-512
317cf8d59f8250129fd699f0d08c86229c05beccf837a1e4b672695d63589c3c820b50d108218fd2a72eb9e3bc96d48e88287fec7c09f3b861d1edab5aca59ac
Type fulltextMimetype application/pdf

Search in DiVA

By author/editor
Matkaselkä, William
By organisation
Department of Computer Science
Computer and Information Sciences

Search outside of DiVA

GoogleGoogle Scholar
Total: 85 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 91 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf