Oculora

A unified ophthalmic VLM for understanding and clinical generation

An ophthalmologist reviewing slit-lamp, fundus, OCT, FFA and B-scan images alongside a clinical question Base Qwen2.5-VL · Params 8.29B · Precision bfloat16 · License Apache-2.0

1.33 M image–text pairs · 360 K ophthalmic instructions · 7 imaging modalities · 300+ ophthalmic conditions

⚠️ Research use only. Oculora is a research artifact — not a medical device. It is not cleared or approved by any regulatory body and must not be used for diagnosis, triage or treatment decisions for real patients. See Intended use and limitations.

Highlights

  • Ophthalmology specialist — Qwen2.5-VL-7B adapted through three stages: pretraining on 1.33 M image–text pairs, supervised fine-tuning on 360 K ophthalmic instructions, and risk-aware DAPO policy optimization.
  • Two evaluation axes — multimodal understanding (multiple-choice and short-answer questions) and open-ended clinical generation (report generation and medical consultation).
  • Broad coverage — a 22,000-sample internal test set plus four external benchmarks spanning seven imaging modalities and over 300 ophthalmic conditions.
  • Leads every free-text benchmark — first on every generation metric across OphthalVQA, FFA-IR and JAMA Challenge against Qwen3-VL, MedGemma and Lingshu (all P < 0.001).
  • Clinically reviewed — ranked highest overall for correctness and control of unsupported content in a 50-case blinded review, with 0.94 agreement (95% CI 0.90–0.97) between LLM-judge and ophthalmologist rankings.
  • Reproducible out of the box — the shipped generation_config.json and preprocessor_config.json reproduce the evaluated configuration exactly.
Full abstract

Previous ophthalmic vision–language model (VLM) studies have commonly evaluated visual question answering and report generation, with limited evidence for medical consultation. Medical consultation requires models to synthesize clinical context and imaging findings into diagnostic and management-oriented responses without predefined answers or explicit diagnostic cues. Here we developed Oculora, a 7B-scale ophthalmic VLM trained on 1.33 million image–text pairs and 360,000 ophthalmic instructions through staged pretraining, supervised fine-tuning and risk-aware policy optimization. We assessed multimodal understanding with multiple-choice and short-answer questions and open-ended clinical generation with report generation and medical consultation on a 22,000-sample internal test set and four benchmarks spanning seven modalities and over 300 ophthalmic conditions. Evaluation combined prespecified task-specific metrics with five-point clinical-quality ratings from large language model judges and senior ophthalmologists. Multiple-choice accuracy was 0.820 internally and 0.674 on JSIEC. Oculora led all generation metrics across three external free-text benchmarks (all P < 0.001). On JAMA Challenge, ROUGE-L was 0.500 versus 0.311 and BERTScore-F1 was 0.927 versus 0.893 for Qwen3-VL-8B. In a 50-case blinded review, Oculora ranked highest overall for correctness and control of unsupported content; LLM and ophthalmologist rankings agreed (0.94, 95% CI 0.90–0.97). These findings show that a specialized ophthalmic VLM can support both multimodal understanding and open-ended clinical generation.

Quickstart

Requires transformers >= 4.57.1. Greedy decoding and max_pixels = 262144 are already the shipped defaults, so loading the model as-is reproduces the reported configuration — no extra flags required.

import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor

model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    "pxiaoyu/Oculora",
    dtype=torch.bfloat16,
    device_map="auto",
)

# max_pixels = 262144 is already the shipped default
processor = AutoProcessor.from_pretrained("pxiaoyu/Oculora")

messages = [{
    "role": "user",
    "content": [
        {"type": "image", "url": "https://example.com/fundus.jpg"},
        {"type": "text", "text": "Describe the fundus findings and give the most likely diagnosis."},
    ],
}]

inputs = processor.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

# greedy decoding is already the shipped default
out = model.generate(**inputs, max_new_tokens=1000)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
vLLM — OpenAI-compatible server
python -m vllm.entrypoints.openai.api_server \
  --model pxiaoyu/Oculora \
  --served-model-name oculora \
  --host 127.0.0.1 --port 8002 \
  --dtype bfloat16 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.75 \
  --max-num-seqs 8 \
  --limit-mm-per-prompt '{"image":4}' \
  --trust-remote-code

Query it as any chat-completions endpoint with image_url content parts. vLLM reads generation_config.json from the repo, so "temperature": 0 is already the default.

Hardware — bfloat16 weights are ~17 GB plus KV cache; a single 40 GB A100 or 48 GB L40S is comfortable at --max-model-len 8192.

Evaluation

Multiple-choice and short-answer questions assess multimodal understanding; report generation and medical consultation assess open-ended clinical generation. Every model below was evaluated under an identical protocol on the same test sets.

Multimodal understanding Oculora Qwen3-VL MedGemma Lingshu
Internal MCQ · Accuracy 0.8197 0.3230 0.3470 0.2993
Internal MCQ · F1 0.8303 0.3393 0.3505 0.3020
JSIEC · Accuracy 0.6740 0.1610 0.1410 0.0650
JSIEC · F1 0.7259 0.2676 0.2175 0.1651
Internal short answer · Unigram F1 0.4869 0.2053 0.1501 0.2114
Internal short answer · ROUGE-L 0.4376 0.1758 0.1120 0.1862
Internal short answer · METEOR 0.4223 0.1466 0.1550 0.1695
Internal short answer · BERTScore-F1 0.8352 0.7593 0.7116 0.7179
OphthalVQA · Unigram F1 0.3189 0.1487 0.1451 0.0664
OphthalVQA · ROUGE-L 0.2749 0.1051 0.1016 0.0560
OphthalVQA · METEOR 0.2719 0.1818 0.1637 0.0434
OphthalVQA · BERTScore-F1 0.7968 0.6945 0.7072 0.6741
Open-ended clinical generation Oculora Qwen3-VL MedGemma Lingshu
FFA-IR · BLEU-2 0.1485 0.0917 0.0579 0.0110
FFA-IR · ROUGE-L 0.2222 0.1639 0.1040 0.0829
FFA-IR · METEOR 0.2393 0.1742 0.1962 0.0581
FFA-IR · BERTScore-F1 0.8870 0.8673 0.8324 0.8492
JAMA Challenge · BLEU-2 0.4466 0.2972 0.0703 0.0183
JAMA Challenge · ROUGE-L 0.5002 0.3106 0.1072 0.0305
JAMA Challenge · METEOR 0.5099 0.3613 0.2930 0.0422
JAMA Challenge · BERTScore-F1 0.9265 0.8925 0.8563 0.8194

Best score per row in bold. Test sets: JSIEC n = 1,000 · OphthalVQA n = 600 · FFA-IR n = 766 · JAMA Challenge n = 177, plus a 22,000-sample internal test set. Baselines are open-weight vision–language models at a comparable parameter scale, run under an identical protocol on the same test sets: Qwen3-VL-8B, MedGemma-4B and Lingshu-7B. Bootstrap 95% confidence intervals for every value are reported in the paper.

Blinded clinical review

Design 50 cases, blinded, ranked by LLM judges and senior ophthalmologists
Outcome Oculora ranked highest overall for correctness and for control of unsupported content
Rater agreement 0.94 (95% CI 0.90 – 0.97) between LLM and ophthalmologist rankings

Training

Three stages, all on the Qwen2.5-VL-7B backbone.

Stage Data Configuration
1 · Pretraining (PT) 1.33 M image–text pairs LoRA, rank 64, α = 128
2 · Supervised fine-tuning (SFT) 360 K ophthalmic instructions LoRA, rank 64, α = 128; vision encoder and multimodal projector frozen
3 · Risk-aware policy optimization — Risk-aware DAPO objective in the VERL framework

Risk-aware DAPO

Setting Value
Responses sampled per prompt n = 8
Global mini-batch per update 128 prompts → 1,024 responses
Risk masking top-K highest-risk generations (K = 0.1) masked globally across the whole mini-batch, before group-relative advantage computation
Reward — keyword 0.5
Reward — LLM-assessed 0.4
Reward — format 0.1

Masking the riskiest decile globally rather than within each group is what stabilizes the gradients: the advantage baseline is never computed over generations the risk model has already flagged.

Inference settings used for all reported results

Setting Value
Decoding greedy
Temperature 0
Max image pixels 262,144

These are the shipped defaults. preprocessor_config.json sets max_pixels = 262144 and generation_config.json sets do_sample = false, so loading the model as-is reproduces the reported configuration — no extra flags required.

Architecture

Class Qwen2_5_VLForConditionalGeneration
Total parameters 8.29 B
Language model 28 layers · hidden 3584 · 28 heads · 4 KV heads (GQA) · SwiGLU · RMSNorm
Vision encoder 32 layers · hidden 1280 → 3584 · 16 heads · patch 14 · spatial merge 2 · window 112
Vocabulary 152,064
Context 128 K positions · sliding window 32,768
Position encoding M-RoPE, mrope_section = [16, 24, 24], θ = 1,000,000
Precision bfloat16
Chat template ChatML (<|im_start|> / <|im_end|>)
Weights 6 safetensors shards · 16.6 GB
Requires transformers >= 4.57.1

Intended use and limitations

Oculora is a research artifact, intended for research on ophthalmic vision–language modelling and for retrospective evaluation on de-identified data.

  • Not a medical device. Not cleared or approved by any regulatory body. It must not be used for diagnosis, triage or treatment decisions for real patients.
  • Outputs can be fluent and confident while being wrong. Every clinical statement requires review by a qualified ophthalmologist.
  • Performance is characterised on the benchmarks above — seven imaging modalities and 300+ conditions. Behaviour outside that distribution (other modalities, populations, languages, or image quality) is unmeasured.
  • Training data inherits the demographic and disease-prevalence biases of its sources; the model may underperform on rare conditions and underrepresented groups.
  • Do not send protected health information to any hosted inference endpoint.

License

Apache 2.0, inheriting the license of the Qwen2.5-VL base model. Users remain responsible for complying with the terms of the underlying base model and of every dataset used for evaluation.

Citation

The paper describing Oculora is currently under review; a full citation will be added here upon publication. Until then, please cite this repository: pxiaoyu/Oculora.

Downloads last month
46
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pxiaoyu/Oculora

Finetuned
(1240)
this model