PP-OCRv6_manga
PP-OCRv6_manga is a lightweight pipeline for manga and anime text detection and recognition (Japanese and Chinese), small enough to run on a CPU: detection 0.9 MB + recognition 10 MB in FP16.
What's new in v0.2 (2026-09-28)
New detector manga_det_v0.2 and recognizer manga_rec_v0.2, drop-in replacements for v0.1 with the same architectures, sizes and speed.
Detector. Retrained on much more manga, webtoon and illustration data. It misses fewer lines, splits or merges lines less often and gives a third fewer false positives on illustrations. With the same recognizer, end-to-end CER drops by about 11% (Table 1).
Recognizer. v0.1 had forgotten most rare kanji; v0.2 brings CER on rare characters from about 60% to about 10% and on SFX from about 45% to about 26%. The other benchmarks improve by about 20% relative CER (Table 2).
Full pipeline, end-to-end CER on the Table 1 pages:
| v0.1 | v0.2 | Change | |
|---|---|---|---|
| Manga | 12.61% | 10.03% | β20.5% |
| Webtoons | 8.38% | 6.21% | β25.9% |
| Illustrations | 18.55% | 16.36% | β11.8% |
| Total | 13.81% | 11.34% | β17.9% |
Models
| Component | Architecture | Paddle | ONNX FP32 | ONNX FP16 |
|---|---|---|---|---|
Det manga_det_v0.2 (default) |
PP-OCRv6 tiny (PPLCNetV4 + RepLKFPN + DBHead) |
1.9 MB | 1.73 MB | 0.92 MB |
Det manga_det_v0.1 |
same | 2.05 MB | 1.73 MB | 0.92 MB |
Rec manga_rec_v0.2 (default) |
PP-OCRv6 small (PPLCNetV4 + lightSVTR + MultiHead CTC/NRTR) |
~21.3 MB | 21.17 MB | 10.63 MB |
Rec manga_rec_v0.1 |
same | ~21.3 MB | 21.17 MB | 10.63 MB |
- Dict:
ppocrv6_dict.txt(18,708 symbols), shared by both recognizers. - The
.pdparamsfiles contain the inference weights (backbone + CTC branch).
Benchmark Results
Evaluation data
- Table 1 (842 pages): the Table 2 pages plus 286 new AnimeText pages; pages found in any training set are excluded.
- Table 2 (562 pages): manga, manhua / webtoon (~85% Chinese) and illustrations; crops from
manga_det_v0.1, corrected ground truth, Korean excluded. - SFX: Manga109 COO onomatopoeia on held-out pages + SFX lines of the Table 2 illustrations.
- Rare kanji: synthetic lines of rare characters in unseen fonts + real AnimeText lines with rare characters.
- JMangaBench_Mixed (3,286 balloon crops, official scoring) and comictxt (217 lines, mean per-line CER).
Table 1: Text Detection Benchmark (recall, false positives, end-to-end CER)
| v0.2 960 |
v0.2 1120 |
v0.1 960 |
v0.1 1120 |
base small 960 |
base medium 960 |
|
|---|---|---|---|---|---|---|
| Size, MB | 1.9 | 1.9 | 2.05 | 2.05 | 10.3 | 64.3 |
| Manga | ||||||
| Recall, % | 94.65 | 94.81 | 94.04 | 94.43 | 90.69 | 92.50 |
| FP | 490 | 467 | 460 | 460 | 466 | 413 |
| CER, % | 10.03 | 9.85 | 11.49 | 11.57 | 13.82 | 10.88 |
| Webtoons | ||||||
| Recall, % | 96.13 | 96.70 | 96.28 | 96.13 | 93.27 | 94.56 |
| FP | 96 | 93 | 108 | 140 | 75 | 49 |
| CER, % | 6.21 | 6.02 | 6.74 | 6.69 | 5.83 | 6.79 |
| Illustrations | ||||||
| Recall, % | 90.02 | 90.68 | 88.94 | 90.26 | 87.33 | 89.84 |
| FP | 297 | 283 | 446 | 344 | 330 | 346 |
| CER, % | 16.36 | 15.04 | 17.79 | 16.25 | 19.03 | 15.90 |
| Total CER, % | 11.34 | 10.89 | 12.75 | 12.46 | 14.70 | 11.87 |
v0.2 / v0.1: manga_det_v0.2 / manga_det_v0.1; base small / medium: the official PP-OCRv6 detectors. 960 / 1120: detector input size; the demo app uses 960. FP: false positives; CER: end-to-end CER. Best in bold, second best underlined.
All detectors run the demo pipeline: thresh=0.15, box_thresh=0.25, unclip_ratio=1.4, area resize for long strips, furigana filter, then manga_rec_v0.2.
Table 2: Text Recognition (CER, lower is better)
| v0.2 | v0.1 | hayai v2.5 nova (p512) |
base medium | base small | |
|---|---|---|---|---|---|
| Size (FP32), MB | 21 | 21 | 599 | 73 | 21 |
| Manga | 5.63 | 6.63 | 7.02 | 9.07 | 10.04 |
| Manhua / webtoon | 4.43 | 5.41 | 13.03 | 4.48 | 5.10 |
| Illustrations | 6.49 | 7.95 | 9.71 | 11.67 | 12.89 |
| SFX (weighted) | 26.14 | 45.03 | 17.01 | 73.21 | 80.71 |
| Rare kanji (weighted) | 10.87 | 60.93 | 80.95ΒΉ | 21.76 | 26.38 |
| JMangaBench | 4.06 | 5.62 | 3.11Β² | 7.19 | 8.38 |
| JMangaBench, text only | 2.28 | 3.56 | 1.78Β² | 4.58 | 5.94 |
| comictxt | 8.37 | 11.39 | 8.15 | 18.98 | 20.17 |
| Japanese | 5.78 | 7.03 | 6.88 | 10.28 | 11.62 |
| Chinese | 3.16 | 4.31 | 13.49 | 3.28 | 3.77 |
| Total | 5.72 | 6.80 | 7.90 | 9.27 | 10.26 |
v0.2 / v0.1: manga_rec_v0.2 / manga_rec_v0.1; hayai: hayai-ocr-v2.5-nova; base medium / small: the official PP-OCRv6 recognizers. CER in %, best in bold, second best underlined. Weighted rows are character-weighted averages over the sets listed under "Evaluation data".
ΒΉ hayai reads the synthetic sets poorly: they are random character strings without context. On the real lines with rare characters it scores 6.87% (v0.2: 6.18%).
Β² hayai reads whole balloon crops, its native input. PaddleOCR models run detection + recognition on the same crops.
The v0.1 numbers above were re-measured the way the demo app runs the model. The v0.1 README reported manga 7.94%; that pipeline padded unsorted batches, which hurt CTC decoding.
Table 3: ONNX FP32 vs FP16 (OpenVINO, CPU)
| Task | Model | FP32 | FP16 |
|---|---|---|---|
| Detection (Table 1 total end-to-end CER) | manga_det_v0.2 |
11.37% | 11.39% |
| Recognition (Table 2 total CER) | manga_rec_v0.2 |
5.72% | 5.79% |
| Recognition (Table 2 total CER) | manga_rec_v0.1 |
6.80% | 6.87% |
FP16 halves the size at a cost of ~0.05 pp CER. ONNX FP32 matches Paddle (detection: 11.34% in Paddle; recognition: 1496/1500 identical predictions).
Directory Structure
PP-OCRv6_manga/
βββ README.md
βββ ppocrv6_dict.txt # Character dictionary (18,708 symbols)
βββ examples/
β βββ onnx_openvino.py # ONNX / OpenVINO pipeline
β βββ paddle_inference.py # Paddle models + PaddleOCR post-processing
βββ det/
β βββ manga_det_v0.2.pdparams # Paddle weights (1.9 MB) <- default
β βββ manga_det_v0.2.onnx # ONNX FP32 (1.73 MB)
β βββ manga_det_v0.2_fp16.onnx # ONNX FP16 (0.92 MB)
β βββ manga_det_v0.1.pdparams # previous version
β βββ manga_det_v0.1.onnx
β βββ manga_det_v0.1_fp16.onnx
βββ rec/
βββ manga_rec_v0.2.pdparams # Paddle inference weights (21.3 MB) <- default
βββ manga_rec_v0.2.onnx # ONNX FP32 (21.2 MB)
βββ manga_rec_v0.2_fp16.onnx # ONNX FP16 (10.6 MB)
βββ manga_rec_v0.1.pdparams # previous version
βββ manga_rec_v0.1.onnx
βββ manga_rec_v0.1_fp16.onnx
Quick Start
- ONNX / OpenVINO (CPU, ~11 MB in FP16, no PaddlePaddle):
examples/onnx_openvino.py, the full pipeline (detection, DB post-processing, recognition) for one page.pip install openvino opencv-python numpy pyclipper python examples/onnx_openvino.py page.jpg - Paddle:
examples/paddle_inference.py, model definitions and weights with PaddleOCR'sDBPostProcessandCTCLabelDecode. - Recommended DB post-processing:
thresh=0.15,box_thresh=0.25,unclip_ratio=1.4, detector input long side 960 (pages longer than 2:1 get the pixel budget of a 3:4 page instead).
Model tree for Kellenok/PP-OCRv6_manga
Base model
PaddlePaddle/PP-OCRv6_small_rec