Skip to content
portfolio

// case study

nihao Chinese Reader Experiment

A technical experiment: a Chinese graded reader for Indonesian learners with AI parsing, word segmentation, dictionary linking, and audio-synced reading.

year
2026
type
Web app
stack
Laravel Development

Diagram of the nihao content pipeline: AI parsing, segmentation, dictionary linking, audio sync, and the reader

nihao was an experiment in building a Chinese graded reader for Indonesian learners. The question I wanted to answer was how much of the work behind a graded reader (splitting text, translating it, linking every word to a dictionary, and syncing audio) could be automated, so that an admin pastes in raw Chinese text and gets back an interactive, levelled story. I built it solo between February and May 2026, then retired it. It isn’t online anymore, so this page describes the engineering rather than a live product.

Content Pipeline:

  • An admin pastes raw Chinese text, and a queued job sends it to DeepSeek, which splits it into sentences and translates each one into Indonesian and English
  • Paragraphs are re-assigned by matching the AI output back to the original text, and the parser recovers from truncated AI responses
  • A Python jieba segmenter, run in batches from Laravel, splits sentences into words
  • Each word is linked to a dictionary of about 124,000 entries imported from CC-CEDICT and tagged with its HSK level; lookups prefer common words over rare readings such as surnames
  • Each story gets a difficulty score and an estimated reading time
  • Audio is transcribed with Whisper (the OpenAI API or a local whisper.cpp), and the transcript is aligned to sentences, with a confidence score per match and flags for low-confidence or implausibly fast segments

Reader:

  • Pinyin shown above the characters, switchable between off, all, and smart mode, which shows pinyin only for words at or above the story’s level
  • Tap any word for its pinyin, meaning, HSK level, examples, and audio, and save it to a vocabulary list
  • Per-paragraph translation toggle, adjustable font size, and audio playback with speed control that highlights the sentence being read
  • Spaced-repetition flashcard review with again, hard, good, and easy ratings

Admin & Engineering:

  • Filament admin for stories, sentences, words, the dictionary, series, and categories, including a one-click AI processing action and tools to split, merge, and fix words
  • Built with Laravel 12, Inertia, Vue 3, and Tailwind CSS 4, with Fortify for authentication and two-factor login
  • 70 Pest test files with about 500 test cases, run in GitHub Actions on PHP 8.4 and 8.5 alongside a lint workflow

Tech Stack:

  • Backend: PHP, Laravel 12, Filament 5, MySQL, queues
  • Frontend: Inertia, Vue 3, Tailwind CSS 4
  • Language processing: DeepSeek, jieba (Python), CC-CEDICT, Whisper
  • Testing: Pest, GitHub Actions