⚠️ Mirror / Backup Notice: This repository is an unmodified backup.

For the latest updates, issues, and support, please visit the upstream repository.

Qwen3.5-OCR-JP-2B — GGUF

GGUF quantized versions of ebinan92/Qwen3.5-ocr-jp-2b, a Japanese/English OCR vision-language model based on Qwen3.5-2B.

Converted and quantized using llama.cpp b10488.

Files

File Size Description
qwen35ocr-jp-2b-q4_k_m.gguf 1.5 GB Q4_K_M quantized (recommended)
qwen35ocr-jp-2b-mmproj-f16.gguf 638 MB Multimodal projector — required for both files below
qwen35ocr-jp-2b-f16.gguf 4.6 GB F16 full precision (for re-quantization)

Both qwen35ocr-jp-2b-q4_k_m.gguf and qwen35ocr-jp-2b-mmproj-f16.gguf are required to run inference.

Usage

llama-server

llama-server \
  -m qwen35ocr-jp-2b-q4_k_m.gguf \
  --mmproj qwen35ocr-jp-2b-mmproj-f16.gguf \
  -ngl 99 \
  --image-min-tokens 1024 \
  --host 127.0.0.1 --port 11436 \
  -c 8192 -n 4096

--host and --port can be changed to suit your environment. Use --host 0.0.0.0 to allow access from other machines on your network.

Ollama (Modelfile)

FROM ./qwen35ocr-jp-2b-q4_k_m.gguf
FROM ./qwen35ocr-jp-2b-mmproj-f16.gguf
ollama create qwen35ocr-jp-2b -f Modelfile

Performance

Measured with Q4_K_M + llama-server b10488 on a single Japanese document page (A4 scanned PDF).

GPU VRAM Speed tok/s
NVIDIA RTX A2000 12 GB ~15 sec/page ~98 tok/s
NVIDIA Quadro P2000 4 GB ~100 sec/page ~43 tok/s

Conversion

# 1. Convert HF model to F16 GGUF
python convert_hf_to_gguf.py ebinan92/Qwen3.5-ocr-jp-2b --outtype f16

# 2. Quantize to Q4_K_M
llama-quantize qwen35ocr-jp-2b-f16.gguf qwen35ocr-jp-2b-q4_k_m.gguf Q4_K_M

Quantization took ~47 seconds on RTX A2000.

About the Original Model

ebinan92/Qwen3.5-ocr-jp-2b is a Japanese/English OCR model fine-tuned from Qwen/Qwen3.5-2B-Base. Key features:

  • Output format: Chandra OCR 2 compatible HTML layout blocks with bounding boxes and semantic labels
  • Japanese handwriting support
  • Vertical text support
  • Ruby (furigana) annotation support (HTML5 <ruby> markup)

Acknowledgements

License

Apache 2.0 — same as the original model.

Downloads last month
289
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for OsGo/Qwen3.5-ocr-jp-2b-GGUF

Finetuned
Qwen/Qwen3.5-2B
Quantized
(2)
this model