Research

Work carried out as a full-time research assistant at the Multilingual Intelligent Mind Service Lab (Mi2S), NCKU. The first item is published. The others are funded or collaborative projects, so they are described here in outline only.

DRPF: evaluating deployment risk in speech recognitionPublished

When older adults answer advance care planning questions by voice, a misrecognized short answer can reverse their meaning while barely moving the character error rate. DRPF sorts recognition errors into three types according to the follow-up each one needs (filter the input, ask the speaker to repeat, or verify by hand) and adds a probe for negated answers rewritten into medical action words. On the same recordings, a cloud pipeline and an edge pipeline with nearly equal short-answer error counts turned out to need different handling.

VOICE: a voice-interaction platform for advance care planningNSTC project

Advance care planning lets people decide, while they can still speak for themselves, which treatments they would not want later. Interviews in this project suggested that older adults get stuck well before any paperwork: they are reluctant to think about it and unsure how to raise it with family. The app therefore first gauges how ready a person is, following the transtheoretical model of behavior change, and shows content for that stage. Questions are answered by retrieving from an advance care planning knowledge graph before a language model writes the reply, so that answers stay within vetted content and each one lists the graph nodes it used. I lead the technical implementation. The prototype has been tried within the team; testing with older adults is planned for the project's second year.

Mandarin–Taiwanese speech recognition for a hospital service deskITRI × Chi Mei Medical Center

Visitors ask the service desk for departments and directions in Mandarin, Taiwanese, or a mix of both. Department names were the hardest part to recognize, and there was no recorded speech for them. I wrote guidance sentences from the hospital's department and location names, synthesized speech for them to extend the training data, fine-tuned the lab's Whisper-based model, traced errors sentence by sentence to decide what data to add next, and deployed the system for field validation at the service desk.

An on-premise document retrieval systemIndustry collaboration

Employees ask questions about internal procedures, memos, and equipment manuals in plain language and check the source page of each answer. Documents and queries may not leave the company network, many files are scanned PDFs, and people search both by describing a problem and by exact part numbers. The system combines vector search with keyword search, runs OCR on scanned pages, and uses a language model hosted inside the company. It was built by a team; my part was the retrieval system and the deployment on the company's own hardware.

iSeeME: an app for guided emotional disclosure and assessmentNSTC project

The user watches a guided video on a tablet and answers its questions aloud while the app records video and audio. The recording is transcribed by the lab's speech recognizer, then scored by a semantic model from a labmate's thesis and a facial-expression model from a partner team, and the scores are combined into a report. I built the app in Flutter, the audio preprocessing that removes silence, background music and the video's own narration before recognition, and the integration of the three models. The app is in development and testing; it was exhibited at a smart-healthcare industry forum in September 2026.