PP-OCRv6_manga

PP-OCRv6_manga is a lightweight pipeline for manga and anime text detection and recognition (Japanese and Chinese), small enough to run on a CPU: detection 0.9 MB + recognition 10 MB in FP16.

What's new in v0.2 (2026-09-28)

New detector manga_det_v0.2 and recognizer manga_rec_v0.2, drop-in replacements for v0.1 with the same architectures, sizes and speed.

Detector. Retrained on much more manga, webtoon and illustration data. It misses fewer lines, splits or merges lines less often and gives a third fewer false positives on illustrations. With the same recognizer, end-to-end CER drops by about 11% (Table 1).

Recognizer. v0.1 had forgotten most rare kanji; v0.2 brings CER on rare characters from about 60% to about 10% and on SFX from about 45% to about 26%. The other benchmarks improve by about 20% relative CER (Table 2).

Full pipeline, end-to-end CER on the Table 1 pages:

v0.1 v0.2 Change
Manga 12.61% 10.03% βˆ’20.5%
Webtoons 8.38% 6.21% βˆ’25.9%
Illustrations 18.55% 16.36% βˆ’11.8%
Total 13.81% 11.34% βˆ’17.9%

Models

Component Architecture Paddle ONNX FP32 ONNX FP16
Det manga_det_v0.2 (default) PP-OCRv6 tiny (PPLCNetV4 + RepLKFPN + DBHead) 1.9 MB 1.73 MB 0.92 MB
Det manga_det_v0.1 same 2.05 MB 1.73 MB 0.92 MB
Rec manga_rec_v0.2 (default) PP-OCRv6 small (PPLCNetV4 + lightSVTR + MultiHead CTC/NRTR) ~21.3 MB 21.17 MB 10.63 MB
Rec manga_rec_v0.1 same ~21.3 MB 21.17 MB 10.63 MB
  • Dict: ppocrv6_dict.txt (18,708 symbols), shared by both recognizers.
  • The .pdparams files contain the inference weights (backbone + CTC branch).

Benchmark Results

Evaluation data

  • Table 1 (842 pages): the Table 2 pages plus 286 new AnimeText pages; pages found in any training set are excluded.
  • Table 2 (562 pages): manga, manhua / webtoon (~85% Chinese) and illustrations; crops from manga_det_v0.1, corrected ground truth, Korean excluded.
  • SFX: Manga109 COO onomatopoeia on held-out pages + SFX lines of the Table 2 illustrations.
  • Rare kanji: synthetic lines of rare characters in unseen fonts + real AnimeText lines with rare characters.
  • JMangaBench_Mixed (3,286 balloon crops, official scoring) and comictxt (217 lines, mean per-line CER).

Table 1: Text Detection Benchmark (recall, false positives, end-to-end CER)

v0.2
960
v0.2
1120
v0.1
960
v0.1
1120
base small
960
base medium
960
Size, MB 1.9 1.9 2.05 2.05 10.3 64.3
Manga
Recall, % 94.65 94.81 94.04 94.43 90.69 92.50
FP 490 467 460 460 466 413
CER, % 10.03 9.85 11.49 11.57 13.82 10.88
Webtoons
Recall, % 96.13 96.70 96.28 96.13 93.27 94.56
FP 96 93 108 140 75 49
CER, % 6.21 6.02 6.74 6.69 5.83 6.79
Illustrations
Recall, % 90.02 90.68 88.94 90.26 87.33 89.84
FP 297 283 446 344 330 346
CER, % 16.36 15.04 17.79 16.25 19.03 15.90
Total CER, % 11.34 10.89 12.75 12.46 14.70 11.87

v0.2 / v0.1: manga_det_v0.2 / manga_det_v0.1; base small / medium: the official PP-OCRv6 detectors. 960 / 1120: detector input size; the demo app uses 960. FP: false positives; CER: end-to-end CER. Best in bold, second best underlined.

All detectors run the demo pipeline: thresh=0.15, box_thresh=0.25, unclip_ratio=1.4, area resize for long strips, furigana filter, then manga_rec_v0.2.


Table 2: Text Recognition (CER, lower is better)

v0.2 v0.1 hayai v2.5 nova
(p512)
base medium base small
Size (FP32), MB 21 21 599 73 21
Manga 5.63 6.63 7.02 9.07 10.04
Manhua / webtoon 4.43 5.41 13.03 4.48 5.10
Illustrations 6.49 7.95 9.71 11.67 12.89
SFX (weighted) 26.14 45.03 17.01 73.21 80.71
Rare kanji (weighted) 10.87 60.93 80.95ΒΉ 21.76 26.38
JMangaBench 4.06 5.62 3.11Β² 7.19 8.38
JMangaBench, text only 2.28 3.56 1.78Β² 4.58 5.94
comictxt 8.37 11.39 8.15 18.98 20.17
Japanese 5.78 7.03 6.88 10.28 11.62
Chinese 3.16 4.31 13.49 3.28 3.77
Total 5.72 6.80 7.90 9.27 10.26

v0.2 / v0.1: manga_rec_v0.2 / manga_rec_v0.1; hayai: hayai-ocr-v2.5-nova; base medium / small: the official PP-OCRv6 recognizers. CER in %, best in bold, second best underlined. Weighted rows are character-weighted averages over the sets listed under "Evaluation data".

ΒΉ hayai reads the synthetic sets poorly: they are random character strings without context. On the real lines with rare characters it scores 6.87% (v0.2: 6.18%).

Β² hayai reads whole balloon crops, its native input. PaddleOCR models run detection + recognition on the same crops.

The v0.1 numbers above were re-measured the way the demo app runs the model. The v0.1 README reported manga 7.94%; that pipeline padded unsorted batches, which hurt CTC decoding.


Table 3: ONNX FP32 vs FP16 (OpenVINO, CPU)

Task Model FP32 FP16
Detection (Table 1 total end-to-end CER) manga_det_v0.2 11.37% 11.39%
Recognition (Table 2 total CER) manga_rec_v0.2 5.72% 5.79%
Recognition (Table 2 total CER) manga_rec_v0.1 6.80% 6.87%

FP16 halves the size at a cost of ~0.05 pp CER. ONNX FP32 matches Paddle (detection: 11.34% in Paddle; recognition: 1496/1500 identical predictions).


Directory Structure

PP-OCRv6_manga/
β”œβ”€β”€ README.md
β”œβ”€β”€ ppocrv6_dict.txt                   # Character dictionary (18,708 symbols)
β”œβ”€β”€ examples/
β”‚   β”œβ”€β”€ onnx_openvino.py               # ONNX / OpenVINO pipeline
β”‚   └── paddle_inference.py            # Paddle models + PaddleOCR post-processing
β”œβ”€β”€ det/
β”‚   β”œβ”€β”€ manga_det_v0.2.pdparams        # Paddle weights (1.9 MB)   <- default
β”‚   β”œβ”€β”€ manga_det_v0.2.onnx            # ONNX FP32 (1.73 MB)
β”‚   β”œβ”€β”€ manga_det_v0.2_fp16.onnx       # ONNX FP16 (0.92 MB)
β”‚   β”œβ”€β”€ manga_det_v0.1.pdparams        # previous version
β”‚   β”œβ”€β”€ manga_det_v0.1.onnx
β”‚   └── manga_det_v0.1_fp16.onnx
└── rec/
    β”œβ”€β”€ manga_rec_v0.2.pdparams        # Paddle inference weights (21.3 MB)   <- default
    β”œβ”€β”€ manga_rec_v0.2.onnx            # ONNX FP32 (21.2 MB)
    β”œβ”€β”€ manga_rec_v0.2_fp16.onnx       # ONNX FP16 (10.6 MB)
    β”œβ”€β”€ manga_rec_v0.1.pdparams        # previous version
    β”œβ”€β”€ manga_rec_v0.1.onnx
    └── manga_rec_v0.1_fp16.onnx

Quick Start

  • ONNX / OpenVINO (CPU, ~11 MB in FP16, no PaddlePaddle): examples/onnx_openvino.py, the full pipeline (detection, DB post-processing, recognition) for one page.
    pip install openvino opencv-python numpy pyclipper
    python examples/onnx_openvino.py page.jpg
    
  • Paddle: examples/paddle_inference.py, model definitions and weights with PaddleOCR's DBPostProcess and CTCLabelDecode.
  • Recommended DB post-processing: thresh=0.15, box_thresh=0.25, unclip_ratio=1.4, detector input long side 960 (pages longer than 2:1 get the pixel budget of a 3:4 page instead).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Kellenok/PP-OCRv6_manga

Finetuned
(3)
this model

Datasets used to train Kellenok/PP-OCRv6_manga

Space using Kellenok/PP-OCRv6_manga 1