🌐 English edition translated by Claude Opus 5 (High reasoning effort). The Traditional Chinese text is authoritative; where the two differ, the Chinese version governs.
Benchmarking Open-Source VLMs on Traditional Chinese OCR and Document Parsing
Only the prompt changed — so why did the very same model score so differently?
That surprise became the question our nine-week VLM OCR benchmarking internship actually set out to study. We started from two independently built evaluation pipelines that produced different scores for the same model, and used OFAT ablation experiments to compare, one factor at a time, how the prompt, the inference parameters, the post-processing and the scorer affected the results. Once the settings were settled, we used TC-STR and OmniDocBench to compare how different VLMs perform on Traditional Chinese short-text recognition and full-page document parsing, recording the models, data, prompts, parameters and raw outputs in full so that the whole pipeline can be re-run and traced.
The experiments showed that a benchmark score does not depend on model capability alone — the evaluation setup itself can change the result substantially: on the same model and the same TC-STR data, Exact Match ranged from 0% to 64.84% across different settings, and adjusting repeat_penalty alone accounted for a difference of 17.78 percentage points. Model size is no direct predictor of performance either: smaller models sometimes beat larger ones, and a given model’s advantage shifts from task to task across short-text recognition, tables, formulas and full-page document parsing.
What we ultimately took away is that a benchmark is not just running models over a dataset once and ranking them by score; it is the combined outcome of the model, the prompt, the inference parameters, the post-processing and the scorer. If those conditions are not controlled, recorded and made public, the scores are not necessarily comparable at all.
So rather than only asking which model scores highest, what matters more is first establishing what we actually measured, and whether that result can be obtained again under consistent, transparent and reproducible conditions.
What counts as evaluating a VLM OCR model “fairly”?
This study belongs to the research track of an OCF-sponsored programme. What we wanted to know was this: when putting an LLM (large language model) to the test, how should we control the many details involved — the prompt, the parameters, the post-processing pipeline, the scorer — so that results from different models can be compared fairly, and can be reproduced and verified by the researchers who come after us?
FLAGSHIP FINDING
Same model, same dataset — change only the evaluation settings, and the results diverge dramatically.
“One small parameter change lifted the model’s score by 17 points.” “A different prompt made the model produce completely different answers.” Situations like these kept coming up throughout our research. As the saying goes, pull one hair and the whole body moves — and evaluating a model is no different. There are so many details and settings involved that any single change can have a sizeable effect on the result. So we used ablation experiments to take these variables apart one at a time, and turned what we learned into an evaluation pipeline that is reproducible, traceable and safe to experiment with, so that comparisons between models rest as far as possible on consistent and transparent conditions.
WHY IT MATTERS
Open data such as government gazettes, court judgments and historical newspapers often exists only as scanned images. If locally runnable open-source VLMs (vision-language models) can be evaluated reliably, they cut both external API (application programming interface) costs and the risk of data leaving the organisation.
RESEARCH QUESTION
Once the task goes beyond ordinary short-text recognition to include table decomposition and document parsing, there is no longer one single way to score a VLM OCR (optical character recognition) answer. Punctuation, whitespace, post-processing and unsolicited explanations can all affect how fair the score is, so we first had to define how to measure.
PROJECT VALUE
The main outcome is evaluation infrastructure that can be re-run, traced back to its sources, keeps failures visible, and can be handed over to the next researcher — so that when we come to build a model of Taiwan’s own, we know how to put it fairly to the test.
02 / TEAM & JOURNEY
Team & Journey
Early in the internship, to get familiar with the evaluation tools and workflow, the two interns each ran their own VLM evaluations. They found that even with the same model and the same benchmark, small differences in prompt, scoring method and parameter settings produced results that diverged widely. After careful ablation experiments and analysis of where the score gap came from, they settled the technical choices used for the formal experiments that followed.
EVALUATION METHOD
游聿堂
Took part in the early Ollama VLM evaluation pipeline and testing
Produced and edited the interactive results report
Reviewed and validated the evaluation pipeline
REPRODUCIBILITY & INFRA
黃以信
Designed the post-processing rules and the four metrics: EM (Exact Match) / CM (Containment Match) / ANLS (Average Normalized Levenshtein Similarity) / F1 (Character F1)
Completed the initial 7 OFAT (One-Factor-At-a-Time) ablation variants and their analysis
Built the long-running evaluation systems benchmark_suite/ and TC-STR_and_OminDoc/
Brought model digest, dataset fingerprint, prompt hash and evaluator commit into provenance tracking
PROJECT TIMELINE · FOUR STAGES
STAGE 01
Scope the question
Settled the research direction and got familiar with the relevant tooling.
STAGE 02
Initial testing
Ablation experiments fixed the technical details for the implementation that followed.
STAGE 03
Full evaluation
Debugged the runs and discussed the evaluation results.
STAGE 04
Report production
Wrote the report and its visualisations, and proofread the content for accuracy.
03 / ABLATION EXPERIMENT
Ablation study: where does the score gap come from?
The two independent pipelines eval/ and Sixhuang/ computed different scores for the same model on the same data, mixing "model capability" together with "evaluation-pipeline effects" and making it hard to say what a score actually meant. To separate the two, we aligned both pipelines onto a single shared baseline and then used an OFAT ablation study (one setting changed per run) to measure each factor's influence in turn.
GLM-OCR Q8_0
One model × 3,706 TC-STR test samples
7
variants — 5 required fresh inference, 2 were recomputed offline
18,530
Ollama requests, with 0 API errors
0% – 64.84%
Range over which EM moves with evaluation settings, same model and data
Each variant changes exactly one setting relative to the baseline. Full definitions of the four metrics — EM (exact match), CM (whether the ground truth is contained in the output), ANLS (normalised similarity based on edit distance) and F1 (the combined character-level precision and recall score) — are given in section 08.
VARIANT
THE ONE CHANGE VS BASELINE
EM↑
CM↑
ANLS↑
F1↑
MEAN LATENCY
Aligned baseline
/api/generate, long prompt, repeat_penalty = 1.6, num_predict = 80
47.06%
76.52%
54.54%
59.17%
0.602 s
Endpoint: /api/chat
Switch to /api/chat, everything else unchanged
47.06%
76.52%
54.54%
59.17%
0.582 s
Sixhuang short prompt
171-character long prompt replaced by a 39-character short prompt
0.00%
79.06%
0.00%
13.56%
0.594 s
Do not set repeat_penalty = 1.6
Fall back to the Ollama / model default instead of passing 1.6 explicitly
64.84%
82.46%
72.08%
76.62%
0.583 s
num_predict = 25
Output length cap lowered from 80 to 25
46.03%
76.34%
53.29%
65.07%
0.293 s
Sixhuang minimal postprocess
The baseline raw responses re-scored with minimal post-processing, computed offline
0.00%
76.79%
0.00%
6.88%
offline recompute
Sixhuang scorer
Baseline predictions unchanged, scored with the other scorer instead
47.87%
77.63%
55.02%
59.96%
offline recompute
KEY 01
repeat_penalty is the single largest lever
Dropping the explicit repeat_penalty = 1.6 and reverting to the default raised EM by 17.78 percentage points: 779 items flipped from wrong to right and 120 from right to wrong, a net gain of 659 fully correct answers. 1.6 is probably too high for this model.
KEY 02
EM at zero ≠ the text was never recognised
The short prompt drove EM and ANLS to zero, yet CM rose to 79.06% — the highest of any variant. The answer was still there in the output; it was just followed by Markdown, explanations, English commentary and a recitation of the prompt, which broke strict matching.
KEY 03
Post-processing decides whether the score is readable at all
The aligned post-processing modified 3,704 of 3,706 raw responses; feed the same outputs through minimal post-processing and EM drops straight to zero. What the score measures today is the performance of "model plus post-processing" as a system, not the model's raw output.
KEY 04
The endpoint can be struck off the suspect list
All 3,706 processed predictions from /api/generate and /api/chat were character-for-character identical, with no difference across the four metrics. A 0.02 s latency gap is not enough to call one faster.
KEY 05
Change the scorer and the score changes
With predictions completely unchanged, swapping the scorer alone marked 30 more items correct (+0.81 pp). The difference comes from whitespace and punctuation normalisation and from whether CM is judged in one direction or both — a reporting-layer difference, not a stronger model. Teams should align scorers before comparing numbers.
KEY 06
Shorter output: 51% faster, with a hidden cost
Lowering num_predict from 80 to 25 cut mean latency by 51% but also cost 1.03 percentage points of EM: some prompt recitations were truncated into incomplete fragments, which then slipped past post-processing rules written for complete phrases.
Recommended configuration from this study
—The aligned long prompt, num_predict = 80, no explicit repeat_penalty = 1.6, the aligned post-processing, and one scorer agreed across the team. Either /api/generate or /api/chat is fine, but a single evaluation run must stick to one of them.
—This combination yielded EM 64.84%, CM 82.46%, ANLS 72.08% and F1 76.62% — the best result among the combinations tested here, and one of the reasons the later benchmark_suite/ fixed its generation settings the way it did.
Limits on interpretation: this is an OFAT single-factor main-effect analysis, so the deltas in each row cannot simply be added together. It also only describes GLM-OCR Q8, TC-STR and the Ollama environment of the time, and does not necessarily generalise to other models. CM only checks whether the prediction contains the ground truth, so an answer buried inside a mass of wrong content can still score — CM alone should not be used to judge OCR quality. And temperature = 0 does not guarantee character-identical output across Ollama versions, inference backends or hardware.
This is a reproducible evaluation toolkit — anyone re-running it should get the same result — that compares several VLMs on two tasks. Both pipelines use Ollama to load models one after another, feed them the same set of images, and record the scores.
TASK A · TC_STR
Traditional Chinese short-text recognition
Recognise short Traditional Chinese text in an image, much like classic OCR (optical character recognition). The dataset is the official test split of esun-ai/traditional-chinese-text-recogn-dataset.
TASK B · OMNIDOCBENCH
Full-page document parsing
Convert a whole document page — headings, tables, formulas, multi-column layout — into structured Markdown. Our original notes spelled it OminDocBench; the correct name is OmniDocBench.
The core commitment is fair comparison: every model sees exactly the same images, prompts and parameters, and nothing is quietly tuned for a model that is doing badly.
05 / PRINCIPLES
Design principles
R1
Same items, same inputs
Within one evaluation run, every model sees the same images, prompts and parameters.
R2
One image per request
Each request carries a single image; images are never batched together.
R3
Models run one at a time
A model is unloaded before the next is loaded, so GPU (graphics processing unit) memory is never shared between models.
R4
Warm-up is not scored
A freshly loaded model has unstable throughput, so warm-up results are excluded from the score.
R5
Bad output stays in
Wrong answers, blank answers and truncated answers all count as the model's own performance; nothing is cleaned up after the fact.
R6
Raw responses are kept forever
Good or bad, every response is stored so it can be checked later.
Together these ensure that a low score reflects model capability rather than a misjudgement caused by a broken test pipeline.
06 / PIPELINE
The end-to-end pipeline
TC_STR and OmniDocBench follow the same pipeline, with two gates — one automatic, one human — along the way.
01
Preparation
Download the dataset, confirm the model list, and record a SHA-256 (Secure Hash Algorithm 256-bit) fingerprint for every image.
02
Preflight checks
Verify the model version and that the GPU is 100% dedicated to this run.
on failure →
BLOCKED
Flagged and aborted; the run does not continue
03
Smoke test
A first pass over 20 TC_STR items or 20 OmniDocBench pages.
04
Human review of the results
HUMAN GATE
The pipeline does not run all the way through by itself: a person must confirm the smoke results look right before the full evaluation begins.
if wrong ↺
Return to step 02 and check again
05
The operator starts the full run manually
This avoids generating piles of bad results with nobody watching.
06
Per-item / per-page inference
Progress is written to a SQLite checkpoint, so an interrupted run can resume.
07
Official scoring
TC_STR is scored by our own code; OmniDocBench uses the official Docker evaluator.
08
Report generation
Interactive HTML report, CSV and JSON.
Signature matching
If the dataset version, model version, prompt or parameters change in any way, the system treats it as a different evaluation run, and old and new results are never scored together.
07 / ENVIRONMENT
System environment
Cloud host
AWS EC2
Operating system
Ubuntu 26.04
GPU
NVIDIA L40S
≈45GB VRAM video RAM
Model runtime
Ollama 0.31.1
Containerisation
Docker
version-pinned
Rebuildable scratch storage
Lost on shutdown but re-downloadable — the dataset itself, for example.
Persistent storage
Holds checkpoints, raw responses and reports, and survives shutdown.
08 / TC_STR
TC_STR: Traditional Chinese short-text recognition
The data is the official test split of esun-ai/traditional-chinese-text-recogn-dataset, 3,706 items in total. The prompt tells the model to output only the text it sees in the image, and is character-for-character identical across all models. Settings: temperature = 0, num_predict = 80, 180-second timeout.
8.1 The 8 models compared
Each card shows the parameter count and the quantisation format. Models marked MoE (Mixture of Experts) are mixture-of-experts models.
Gemma 4 E2B
4.6B
Q4_0
Gemma 4 E4B
7.5B
Q4_0
Gemma 4 12B
11.9B
Q4_0
Gemma 4 26B A4B
25.2B
≈3.8B active
Q4_0MoE
Gemma 4 31B
30.7B
Q4_0
GLM OCR BF16 (0.9B)
0.9B
F16scoring exception
GLM-4.6V-Flash 9B
9.4B
Q4_K_M
Kimi-VL-A3B-Instruct
16B
≈3B active
Q4_K_MMoE
8.2 Scoring metrics (higher is better for all four)
Exact Match
Counted correct only if the answer matches the ground truth (GT) character for character.
Containment Match
Whether the ground truth appears in full inside the model's answer (extra output is allowed).
ANLS
A similarity derived from edit distance; anything below 0.5 is scored as 0.
Character F1
Compares character by character, combining how much was right with how much extra was produced.
Scoring exception: GLM OCR BF16 (0.9B) sometimes returned perfectly normal text while also reporting done=false and omitting its statistics fields. After manual review, the rule was relaxed for this model only: any returned text is scored, but flagged as "truncation cannot be determined", and its efficiency figures are not comparable with the other models.
09 / OMNIDOCBENCH
OmniDocBench: full-page document parsing
We use opendatalab/OmniDocBench from Hugging Face, pinned to v1.6, 1,651 pages in total, covering multilingual content, multi-column layouts, tables, formulas and blurry scans. The prompt asks for the whole page as Markdown: reading order preserved, tables converted to HTML tables that keep merged cells, and formulas expressed in LaTeX.
Unlike TC_STR, which we score ourselves, this task uses OmniDocBench's official scoring program at a pinned version, run inside a Docker container so the scoring logic cannot drift with the host environment.
↓ lower is better
Text edit distance
How far the plain text is from the ground truth, after normalisation.
↑ higher is better
Formula CDM (Character Detection Matching)
Whether formula structure and content match correctly.
↑ higher is better
Table TEDS (Tree-Edit-Distance-based Similarity)
Similarity of the table's tree structure.
↓ lower is better
Reading Order
Whether reading order matches the ground truth.
official composite
Overall
((1−text edit)×100 + TEDS + CDM) ÷ 3
We also keep the same four supplementary diagnostic metrics used for TC_STR, but label them "not official leaderboard metrics" so they are not confused with the official scores.
DEBUG LOG
Repetition loops hang the official evaluator
Running the full test set with smaller models such as Gemma 4 E2B, roughly 30% of the outputs fall into endless repetition of the answer. Once such output reaches the official matching stage, even the built-in chunked Hungarian algorithm fallback for long inputs hangs and never returns a result.
match_workers: 4 → 1
Lowering the evaluator's worker-process count avoids the hang (upstream issue #228).
temperature: 1.0
Too low a temperature readily triggers repetition, so OmniDocBench is fixed at 1.0 (TC_STR uses 0).
10 / RESULTS
Results
Every bar below is scaled in direct proportion to the value in the corresponding table, so the graphics and the numbers agree exactly.
10.1 Run parameters and settings
Both benchmarks follow the same design: one fixed prompt plus one fixed set of generation parameters, applied to every model. The config files contain no per-model overrides — the only thing that differs between models is the tag, the quantisation and the model id. Every value below is taken directly from the config files as they were actually run.
TC-STRTC_STR/config.yaml · models.yaml
All 8 models share one set of parameters and one prompt: tc_str_bench/config.py requires exactly 8 model ids, and settings.prompt / settings.options are read globally, with no per-model override.
Ollama call
endpoint/api/generate
streamfalse
thinkfalse
keep_alive5m
timeout180 s
max_attempts3
retry_backoff2 s · 5 s
Note: keep_alive = 5m is passed in by the runner at call time; the keep_alive: 0 in config.yaml is used only for model preload/unload.
The prompt was sent to the models in Traditional Chinese and is reproduced verbatim above. In English: “You are a professional OCR engine. Look carefully at the image and output only the text you see in it (mainly Traditional Chinese; also output any English or digits that appear). Return the plain text itself only — no HTML tags, Markdown, JSON or any other formatting markup; no explanation, prefix or suffix; nothing beyond the text and its punctuation; no translation; no quotation marks. Where text is blurred, skewed or occluded, infer the most likely character from its shape.”
OMNIDOCBENCHOminDocBench/config/benchmark.json
All models share one set of inference parameters and one prompt: the report block in core.py labels them “common prompt/options” and records a SHA-256 hash to prove they stayed identical.
Ollama call
endpoint/api/generate
streamfalse
thinkfalse
batch_size1
keep_alive10m
timeout1800 s
max_attempts3
backoff5 s · 20 s
kv_cache_typefp16
Generation options
num_ctx65536
num_predict16384
temperature1.0
repeat_penalty1.0
top_k64
top_p0.95
seed20260731
Prompt · full page to Markdown
請將這一頁文件完整解析為 Markdown。依照原始閱讀順序保留所有可見文字;標題、段落、清單與程式碼請使用適當的 Markdown;表格請以能保留列、欄與合併儲存格結構的 HTML table 輸出;行內公式與獨立公式請使用 LaTeX;不要描述圖片內容,不要加入解釋、前言、結語或 code fence,只輸出解析後的文件內容。
In English: “Parse this page of the document completely into Markdown. Preserve all visible text in the original reading order; use appropriate Markdown for headings, paragraphs, lists and code; output tables as HTML tables that preserve the row, column and merged-cell structure; use LaTeX for both inline and display formulas; do not describe images, and do not add explanation, preamble, closing remarks or code fences — output only the parsed document content.”
prompt_source marks this as a user-specified fallback: the official OmniDocBench v1.6 repo ships neither a Gemma-specific inference script nor a generic end-to-end VLM prompt.
Sampling parameters: identical across both benchmarks
top_k and top_p are hard-coded in both config files rather than inherited from Ollama’s built-in defaults, and both benchmarks use exactly the same values. The real differences are temperature (0 for TC-STR, 1.0 for OmniDocBench, to suppress repetition loops) and num_predict (80 vs 16384).
Parameter
TC-STR
OmniDocBench
top_k
64
64
top_p
0.95
0.95
OmniDocBench scorer parameters (official evaluator, not model-side settings)
—match_workers was lowered from the official default of 4 to 1 to stop repetition-loop output from small models hanging the scorer (see the DEBUG LOG in section 09).
10.2 TC_STR full-set results
The complete 3,706-item test set, ordered by EM from high to low. Values are taken from the primary leaderboard of run tcstr_20260729T162125139376Z.
Why does GLM-4.6V-Flash (9.4B) beat Gemma 4 31B (30.7B)?
Comparing item by item shows that GLM’s advantage is not spread evenly across the test set: it leads by 14.2 percentage points on the sign category, well above its 6.8 pp on billboard and 4.5 pp on poster. Sampling 194 sign items that GLM answered correctly and Gemma got wrong, we found that Gemma frequently misreads visually similar characters in small cropped text, and is also more likely to give up on an answer outright when an image is blurry or too low in resolution. These behaviours may be part of why the gap is widest on the sign category.
On the evidence so far, then, we lean towards the explanation that the two models differ in how well they recognise difficult short text, rather than the gap being a matter of parameter count alone.
POST-HOC CHECK
Gemma 4 26B A4B (MoE, ≈3.8B active) edges past the 31B dense model
In the TC-STR short-text OCR test, Gemma 4 26B A4B — a MoE architecture that activates only about 3.8B parameters per inference — still edges out Gemma 4 31B, 82.0% to 80.6%. On this kind of short-text recognition task, then, a larger model does not necessarily score higher.
On the OmniDocBench full-page parsing task, however, the result is exactly reversed: the 31B leads 26B A4B by 8.5 points. This suggests 26B A4B’s advantage may be tied to the nature of the task, and does not show that a MoE architecture is inherently better than a dense one. (See section 10.3 below.)
EM versus CM
0–100% · higher is better
EM exact matchCM containment match
GLM-4.6V-Flash 9B
88.4
88.5
Gemma 4 26B A4B
82.0
82.5
Gemma 4 31B
80.6
81.2
Gemma 4 12B
58.9
60.4
Gemma 4 E4B
49.0
49.2
Gemma 4 E2B
43.2
43.8
Kimi-VL-A3B-Instruct
38.5
39.6
GLM OCR BF16 (0.9B)
36.7
77.2
GLM OCR BF16 (0.9B) is the only model whose EM falls far below its CM — a gap of 40.5 percentage points.
POST-HOC CHECK
Why is the EM/CM gap for GLM OCR BF16 (0.9B) so wide?
GLM OCR BF16 (0.9B)’s EM and CM differ by 40.5 percentage points, meaning many answers do contain the correct text but fail EM because the model emitted other content alongside it. Unfortunately the per-item raw output for the BF16 run was not retained, so we used data from the same model family — GLM-OCR Q8_0 — as a reference for the analysis.
Among Q8_0’s 3,706 answers, about 19.8% show format breakdown: stray Markdown markers, for instance, or continued generation after the answer was already correct. These cases score 0% EM but 76% CM, the same gap pattern seen with BF16.
That said, genuine cases of the same sentence repeating over and over account for only 0.24%. Far more common is the model getting the answer right and then continuing to emit extra content. The longer the answer, moreover, the higher the rate of format breakdown. Together these results suggest that over-generation may be one of the main reasons for the EM/CM gap.
Mean latency
seconds per item · lower is better
GLM-4.6V-Flash 9B
0.41s
Gemma 4 26B A4B
1.18s
Gemma 4 31B
1.20s
Gemma 4 12B
1.27s
Gemma 4 E4B
1.22s
Gemma 4 E2B
1.16s
Kimi-VL-A3B-Instruct
0.33s
GLM OCR BF16 (0.9B)
0.27s
Share of truncated samples
out of 3,706 items · lower is better
GLM-4.6V-Flash 9B
0.0%
Gemma 4 26B A4B
0.2%
Gemma 4 31B
0.0%
Gemma 4 12B
0.1%
Gemma 4 E4B
2.9%
Gemma 4 E2B
3.4%
Kimi-VL-A3B-Instruct
12.2%
GLM OCR BF16 (0.9B)
54.9%
POST-HOC CHECK
Kimi-VL-A3B’s high truncation rate comes mainly from runaway repetition
Per-item analysis shows that all 453 of Kimi-VL-A3B’s truncated cases hit the 80-token (a token being the smallest unit of text a model processes) generation limit, and 452 of them do not even contain the correct answer. Most outputs fall into a repetition loop over a single character, a short word or the prompt itself, rather than being answers too long to finish.
The problem clusters particularly around 1–4 character short answers and signboard images. On the same items, Gemma 4 26B A4B shows only a handful of comparable cases, which indicates that the difficulty of the data alone is not enough to explain the result.
Both models ran with the same generation settings: temperature=0 and repeat_penalty=1.0. With that combination, once a model starts repeating it has no mechanism for breaking out of the loop, so a brief recognition slip can be amplified into generation that runs all the way to the limit. Why Kimi-VL falls into such loops more readily than Gemma remains unconfirmed, however: the quantisation build, the model’s own generation stability and the template configuration would all need further experiments to tell apart.
#
MODEL
EM↑
CM↑
ANLS↑
F1↑
LATENCY
TRUNCATED
1
GLM-4.6V-Flash 9B
88.37%
88.45%
94.07%
94.30%
0.41s
0 / 3,706
2
Gemma 4 26B A4B
81.98%
82.46%
88.71%
89.58%
1.18s
7 / 3,706
3
Gemma 4 31B
80.63%
81.17%
87.82%
88.55%
1.20s
0 / 3,706
4
Gemma 4 12B
58.90%
60.39%
72.89%
75.06%
1.27s
4 / 3,706
5
Gemma 4 E4B
48.95%
49.19%
69.99%
72.87%
1.22s
109 / 3,706
6
Gemma 4 E2B
43.20%
43.82%
66.52%
70.09%
1.16s
126 / 3,706
7
Kimi-VL-A3B-Instruct
38.51%
39.64%
49.86%
53.69%
0.33s
453 / 3,706
8
GLM OCR BF16 (0.9B)
36.70%
77.23%
40.72%
45.41%
0.27s
2,034 / 3,706
METRIC CAVEAT
The 0.5 ANLS threshold is too coarse for short Chinese terms
Conventional ANLS applies a threshold of 0.5: a similarity of 0.5 or above counts as a valid recognition, anything below it is zeroed out. That works well for alphabetic scripts, where an English word can survive a few misspelt letters. For the terse Chinese terms in this dataset — two- or three-character proper nouns, numbers, codes — a single wrong character is enough to drop the score below 0.5 and have it scored as 0, which is neither fair nor fine-grained enough for Chinese.
KEY 01
GLM-4.6V-Flash 9B is strongest across the board
Highest on all four of EM, CM, ANLS and F1, with zero truncations, zero anomalies and a 100% success rate — and among the fastest at 0.41 s per item.
KEY 02
Bigger Gemma is more accurate, but not linearly
26B A4B (MoE, ≈3.8B active) edges past the larger dense 31B model, 82.0 to 80.6 — unlike section 10.3, where scores rise strictly with parameter count.
KEY 03
The EM/CM gap for GLM OCR BF16 (0.9B)
54.9% of samples (2,034 of 3,706) are flagged as truncated. The answer is often buried inside the output — hence the high CM — but format breakdown loses the points under strict matching, hence the low EM.
Further observations
—Inference speed tracks architecture more directly than parameter count: the three non-Gemma models all sit at 0.27–0.41 s, far faster than every Gemma 4 (1.15–1.27 s). Quantisation format (F16 / Q4_K_M vs Q4_0) or inference architecture may be responsible.
—Kimi-VL-A3B-Instruct is a MoE with few active parameters, yet its anomaly rate is far from low (453 truncations) and its EM only 38.5%. MoE alone guarantees no stability; it still comes down to each model's training and quantisation quality.
10.3 OmniDocBench full-corpus results
Official evaluator scores over the full 1,651-page corpus, ordered by Overall from high to low. Overall draws on just three components: text, table and formula.
Overall score ranking
0–100 · higher is better
Qwen3-VL 32B
85.3
Qwen3-VL 4B
75.6
Gemma 4 31B
71.9
InternVL3.5 38B
70.5
InternVL3.5 4B
69.6
Gemma 4 26B A4B
63.4
Gemma 4 12B
37.5
Gemma 4 E4B
26.9
Gemma 4 E2B
17.3
Qwen3-VLGemma 4InternVL3.5
POST-HOC CHECK
The short-text OCR advantage disappears once the task becomes full-page parsing
In TC_STR short-text OCR, Gemma 4 26B A4B narrowly beat Gemma 4 31B. On OmniDocBench full-page document parsing the result flips completely: 26B A4B scores 63.4 against the 31B’s 71.9, a lead of 8.5 points.
Our conjecture is that this relates to task complexity. Beyond recognising text, full-page document parsing also has to handle layout, tables and formulas at the same time, so it calls for the integration of more capabilities than short-text OCR does.
Capability radar
all rescaled to 0–100; further out is better
Text and Order have been converted from edit distance to accuracy (1 − edit). Select models on the right to overlay them.
Choose models to overlay
Qwen3-VL 32B
Text
88.9
TEDS
78.5
CDM
88.5
Order
79.4
Overall
85.3
Qwen3-VL 4B
Text
85.1
TEDS
67.6
CDM
74.2
Order
75.7
Overall
75.6
Gemma 4 31B
Text
71.4
TEDS
63.2
CDM
81.3
Order
72.4
Overall
71.9
InternVL3.5 38B
Text
74.5
TEDS
57.9
CDM
79.0
Order
71.8
Overall
70.5
InternVL3.5 4B
Text
80.4
TEDS
59.3
CDM
69.3
Order
74.2
Overall
69.6
Gemma 4 26B A4B
Text
62.6
TEDS
55.3
CDM
72.4
Order
64.9
Overall
63.4
Gemma 4 12B
Text
35.6
TEDS
30.1
CDM
46.7
Order
49.8
Overall
37.5
Gemma 4 E4B
Text
28.6
TEDS
20.4
CDM
31.7
Order
45.1
Overall
26.9
Gemma 4 E2B
Text
19.9
TEDS
6.8
CDM
25.2
Order
37.3
Overall
17.3
#
MODEL
OVERALL↑
TEXT EDIT↓
TEDS↑
CDM↑
ORDER↓
1
Qwen3-VL 32B
85.31
0.111
0.785
0.885
0.206
2
Qwen3-VL 4B
75.62
0.149
0.676
0.742
0.243
3
Gemma 4 31B
71.93
0.286
0.632
0.813
0.276
4
InternVL3.5 38B
70.47
0.255
0.579
0.790
0.282
5
InternVL3.5 4B
69.64
0.196
0.593
0.693
0.258
6
Gemma 4 26B A4B
63.43
0.374
0.553
0.724
0.351
7
Gemma 4 12B
37.46
0.644
0.301
0.467
0.502
8
Gemma 4 E4B
26.89
0.714
0.204
0.317
0.549
9
Gemma 4 E2B
17.32
0.801
0.068
0.252
0.627
KEY 01
Qwen3-VL 32B leads across the board
Its Overall opens a near-10-point lead over second place, and it posts the best table TEDS and formula CDM of all nine models.
KEY 02
Gemma 4 scales almost in proportion to parameter count
It climbs steadily from E2B (17.3) all the way to 31B (71.9) with no plateau along the way, showing that document parsing depends heavily on model scale.
KEY 03
Small models do not necessarily lose to large ones
Qwen3-VL 4B (75.6) beats Gemma 4 31B, and scaling InternVL3.5 from 4B (69.6) to 38B (70.5) buys only a marginal gain.
Further observations
—Gemma 4 E2B's table TEDS is only 0.068 — effectively "cannot read tables at all". Scaling the same family up to 31B only reaches 0.632, so table-structure understanding is especially hard for small models.
—Formula CDM is generally the relatively strongest metric across all nine models; even the smallest, E2B, reaches 0.252. We suspect this relates to the comparatively fixed structure of LaTeX notation.
—Both Qwen3-VL 4B and InternVL3.5 4B post a higher Overall than the nominally larger Gemma 4 26B A4B (63.4): architecture and training-data quality sometimes matter more than the parameter count on paper.
POST-HOC CHECK
Gemma 4 E2B's table score collapses to 0.068 — is layout type really the cause?
We initially suspected that complex layouts — tables spanning pages, merged cells, multi-column text — were the cause, but OmniDocBench does not provide complete annotations for these table types, so the idea cannot be verified directly. The available data even shows that pages containing tables are not particularly concentrated in multi-column layouts, which leaves the multi-column explanation without support.
Another possible lead is that small models often fall into long repetitions, while tables generally require longer, structurally more complex output. Our conjecture is that this makes E2B more prone to losing control on table tasks, dragging its score down.
POST-HOC CHECK
Formula recognition is less affected by model scale — but not entirely unaffected
We first thought the formulas in OmniDocBench might simply be easier, which would explain how small models still score reasonably well. But checking all 2,066 formulas showed the test set in fact contains a fair number of matrices and complex LaTeX structures, so “the questions are too easy” is not enough to explain it. The results do show that formula CDM falls off far less sharply than table TEDS as models get smaller. CDM still rises with model scale, though, so the more accurate statement is this: formula recognition is relatively insensitive to model scale, not wholly unaffected by it.
Our conjecture is that this relates to the comparatively fixed visual patterns of the symbols and structures found in formulas. Tables, by contrast, also require understanding column and row positions and relationships across cells, and so may depend more on the model’s overall capability. No per-formula analysis is available to confirm this directly, however.
10.4 How results are presented
The README deliberately avoids hard-coding in-progress scores into the documentation; the authoritative numbers are the reports in each run's output directory. The tooling also produces:
An interactive comparison report (HTML)
Sortable by any metric, with filters for best / worst / random-sample pages, so a model's weak scenarios surface quickly.
Detailed statistics
Success rate, API failure count, retry count, blank-output count and mean latency — separating model-capability problems from systems-engineering problems.
Per-page / per-item raw records
The model's raw answers are retained in full for later investigation or manual review.
11 / ENGINEERING CHALLENGES
The core technical challenge: making a score trustworthy
The three cases below best capture the research-engineering character of this internship. All three are rooted in experimental design and directly affect whether a score can be interpreted at all — they are not merely coding mistakes.
Differences in prompt, answer handling and scoring rules mean a score gap may not come from the model alone.
Action
Align both test pipelines onto the same settings first, then change one condition at a time — 7 configurations compared in total, 5 of which required re-running the models.
Outcome
The results show that changing just one setting can move the EM score substantially. From then on, every experiment recorded in full the settings actually used.
CHALLENGE 02 · OFFICIAL EVALUATOR
Long repetitive model output stalls the OmniDocBench scoring program
→ DEBUG UPSTREAM
Problem
Some small models repeat the same content endlessly, in roughly 30% of cases. These over-long answers not only hurt the model’s own results but can also stop OmniDocBench’s scoring program from completing.
Action
Raising temperature from 0 to 1 markedly reduced long repetitive output. Following upstream issue #228, we also lowered match_workers from 4 to 1, and ran a small-scale test first to confirm the settings executed stably. The models’ raw output is retained in full to make later investigation easier.
Outcome
After the adjustments, the scoring pipeline completes reliably. Repetitive model output is still counted as-is in the results, and is not deleted or corrected before scoring.
CHALLENGE 03 · PROVENANCE
Same model name, different version, third-party quantisation — is it still the same model?
→ TRACE EVERYTHING
Problem
If a model tag, dataset revision or evaluator commit changes quietly, old and new results can end up wrongly mixed together.
Action
The resume key and run signature incorporate model digest, dataset fingerprint, prompt hash and dependency lock; third-party GGUF builds have their source and risks recorded separately.
Outcome
Any difference in a key setting counts as a new experiment and never runs against old data, and every score can be traced back through metadata to the conditions it was actually produced under.
12 / DELIVERABLES & REFLECTION
What the internship leaves behind
01 · EVALUATION PIPELINES
Re-runnable evaluation tooling
eval/, ablation_experiment/, benchmark_suite/ and TC-STR_and_OminDoc/, covering short-text OCR, full-page parsing and multi-model comparison.
02 · REPRODUCIBILITY METADATA
Results that carry proof of provenance
Run manifests, dataset manifests, model digests, prompt hashes, version pinning and protocol exceptions, so that experiments sharing a name but differing in substance are never conflated.
03 · EVIDENCE & REPORTS
Raw evidence and interactive output
Raw responses, per-item CSV / JSON / SQLite, ablation results and self-contained HTML reports, so conclusions can be checked by a human instead of collapsing into one final percentage.
04 · HANDOFF DOCUMENTATION
Handover-ready research documentation
Methodology, metrics, version pins, model provenance, runtime environment and exception records are all written down, reducing the risk that the research exists only in its original authors' memories.
WHAT THIS INTERNSHIP TAUGHT US
A benchmark score is the product of the whole pipeline
Model, prompt, generation options, post-processing and scorer can all move the result. Before comparing models, confirm that the measurement is aligned.
Reproducibility is not an appendix
A long-running benchmark has to be designed from the start with resume, version identity, anomaly logging and raw-output retention — not documented after the run is over.
Failure cases are research data in their own right
Repetition, truncation, missing metadata and evaluator hangs should not be silently cleared away. Keeping them is how you learn where the system's real limits are.
Follow-up post-hoc checks: results for six hypotheses (see the "post-hoc check" blocks below the charts in 10.2 and 10.3)
—Not one of the six hypotheses could be fully confirmed. The root cause is that the per-item and per-page raw outputs of the TC_STR bench and OmniDocBench were never committed to the repository — only the aggregate scores were. Which rather proves the lesson above: reproducibility is not an appendix.
—The most important correction: Gemma 26B A4B's "small beats large" holds only on the TC-STR short-text task. On OmniDocBench full-page parsing it loses to the 31B by 8.5 points. This is a task-dependent one-off, not a durable MoE advantage.
—Kimi-VL's instability was originally blamed on the MoE architecture. Contrary evidence overturned that guess and pointed instead to the better-grounded lead of a third-party, unofficial quantisation.
—GLM-4.6V-Flash beating Gemma 31B is not a scoring artefact — that much is ruled out — but why it wins, at the level of architecture or training data, remains unsolved, pending the per-item data needed to check it.
The core value of this project is a rigorous, reproducible, verifiable evaluation method; which model is stronger is a by-product along the way. The approach: fix the data, the prompts and the generation settings, and pair them with hash fingerprints, version pinning, checkpoints and exception records, so that every score can be traced back to the specific conditions that produced it. This cannot eliminate every benchmark bias, but it does make the non-model factors in this comparison more transparent, and it tells later researchers which results are comparable and which come with caveats attached.
Limitations
GLM OCR BF16 (0.9B) must be read alongside its exception rule Because completion metadata was missing, relaxed scoring was applied, so its EM/CM gap and truncation rate are not on exactly the same footing as the other seven models.
Some models are community conversions InternVL3.5 (4B / 38B) are community-converted GGUF quantisations, whose performance may differ from the official native builds.
MoE "active parameter" figures are estimates The active-parameter counts given for Gemma 4 26B A4B and Kimi-VL-A3B-Instruct are conceptual estimates, not official exact figures.
Sampled reports and full-corpus results are not interchangeable The interactive comparison report takes only 5 best, 5 worst and 5 random pages (15 in total); this page uses the official scores over the full 1,651-page corpus.
Cross-environment reproducibility has limits Even with temperature=0 fixed and digests and dataset versions pinned, a change of Ollama version, driver or hardware still does not guarantee character-identical output.
Future work
Error analysis of GLM OCR BF16 (0.9B)'s anomalous samples Review the 2,034 truncated samples item by item to establish which kinds of text — long sentences, dense small print, unusual symbols — most readily trigger repetition.
Detailed analysis of the weakest sub-metrics For small models' table TEDS, identify which layouts (page-spanning, merged cells, multi-column) are losing the points.
A finer repeat_penalty sweep and interaction testing The ablation in section 03 showed 1.6 to be clearly harmful but did not locate the optimum, and the interactions among prompt, post-processing and num_predict remain untested — OFAT cannot capture those interaction effects.
Consolidation into a continuously updated dashboard Merge the interactive comparison report with the static charts, so a newly evaluated model folds into the comparison automatically.
14 / ACKNOWLEDGEMENTS
Acknowledgements
This report was funded by the FreeSEED surplus. Our thanks go to everyone who donated to FreeSEED.
We thank Huang Liang-Hsun (黃亮勳), founder of Twinkle AI, for his technical guidance.
This English edition was translated from the Traditional Chinese original by Claude Opus 5 (High reasoning effort). Numbers, model names, identifiers and code paths are carried over verbatim; where any wording diverges, the Traditional Chinese version is authoritative.