Reports · 報告列表
TR-00
開源 VLM 繁體中文 OCR 與文件解析評測
Benchmarking Open-Source VLMs on Traditional Chinese OCR and Document Parsing
九週實習成果報告:比較兩套獨立建立的評測流程,並以消融實驗拆解 prompt、推論參數、後處理與評分器對分數的影響,最後收斂成一套可重現、可追溯的評測流程。
A nine-week internship report: two independently built evaluation pipelines compared side by side, with ablation studies isolating how prompts, inference parameters, post-processing, and scorers move the numbers — and the reproducible, traceable evaluation workflow that came out of it.
TR-01
粵文 OCR 評測資料集
Cantonese OCR Benchmark
合成影像評測資料集:涵蓋單字、詞彙與句子三個層級共 4,930 張影像,並聚焦 154 個罕見於標準中文的粵文專用字。以兩個 4B 級視覺語言模型(Qwen3-VL、InternVL3.5)實測,指出擴充區 Unicode 字元的辨識失效,並區分視覺辨識不足與語言模型補償兩種效應。
A synthetic image benchmark for Cantonese OCR: 4,930 images across character, vocabulary, and sentence levels, focusing on 154 Cantonese-specific characters that are rare in standard Chinese. Two 4B vision-language models (Qwen3-VL and InternVL3.5) are evaluated, exposing recognition failures on extended-Unicode characters and separating visual recognition deficits from language-model compensation.