Instructions to use pxiaoyu/Oculora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pxiaoyu/Oculora with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="pxiaoyu/Oculora") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("pxiaoyu/Oculora") model = AutoModelForMultimodalLM.from_pretrained("pxiaoyu/Oculora", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use pxiaoyu/Oculora with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pxiaoyu/Oculora" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pxiaoyu/Oculora", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/pxiaoyu/Oculora
- SGLang
How to use pxiaoyu/Oculora with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pxiaoyu/Oculora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pxiaoyu/Oculora", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pxiaoyu/Oculora" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pxiaoyu/Oculora", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use pxiaoyu/Oculora with Docker Model Runner:
docker model run hf.co/pxiaoyu/Oculora
Oculora
A unified ophthalmic VLM for understanding and clinical generation
1.33 M image–text pairs · 360 K ophthalmic instructions · 7 imaging modalities · 300+ ophthalmic conditions
⚠️ Research use only. Oculora is a research artifact — not a medical device. It is not cleared or approved by any regulatory body and must not be used for diagnosis, triage or treatment decisions for real patients. See Intended use and limitations.
Highlights
- Ophthalmology specialist — Qwen2.5-VL-7B adapted through three stages: pretraining on 1.33 M image–text pairs, supervised fine-tuning on 360 K ophthalmic instructions, and risk-aware DAPO policy optimization.
- Two evaluation axes — multimodal understanding (multiple-choice and short-answer questions) and open-ended clinical generation (report generation and medical consultation).
- Broad coverage — a 22,000-sample internal test set plus four external benchmarks spanning seven imaging modalities and over 300 ophthalmic conditions.
- Leads every free-text benchmark — first on every generation metric across OphthalVQA, FFA-IR and JAMA Challenge against Qwen3-VL, MedGemma and Lingshu (all P < 0.001).
- Clinically reviewed — ranked highest overall for correctness and control of unsupported content in a 50-case blinded review, with 0.94 agreement (95% CI 0.90–0.97) between LLM-judge and ophthalmologist rankings.
- Reproducible out of the box — the shipped
generation_config.jsonandpreprocessor_config.jsonreproduce the evaluated configuration exactly.
Full abstract
Previous ophthalmic vision–language model (VLM) studies have commonly evaluated visual question answering and report generation, with limited evidence for medical consultation. Medical consultation requires models to synthesize clinical context and imaging findings into diagnostic and management-oriented responses without predefined answers or explicit diagnostic cues. Here we developed Oculora, a 7B-scale ophthalmic VLM trained on 1.33 million image–text pairs and 360,000 ophthalmic instructions through staged pretraining, supervised fine-tuning and risk-aware policy optimization. We assessed multimodal understanding with multiple-choice and short-answer questions and open-ended clinical generation with report generation and medical consultation on a 22,000-sample internal test set and four benchmarks spanning seven modalities and over 300 ophthalmic conditions. Evaluation combined prespecified task-specific metrics with five-point clinical-quality ratings from large language model judges and senior ophthalmologists. Multiple-choice accuracy was 0.820 internally and 0.674 on JSIEC. Oculora led all generation metrics across three external free-text benchmarks (all P < 0.001). On JAMA Challenge, ROUGE-L was 0.500 versus 0.311 and BERTScore-F1 was 0.927 versus 0.893 for Qwen3-VL-8B. In a 50-case blinded review, Oculora ranked highest overall for correctness and control of unsupported content; LLM and ophthalmologist rankings agreed (0.94, 95% CI 0.90–0.97). These findings show that a specialized ophthalmic VLM can support both multimodal understanding and open-ended clinical generation.
Quickstart
Requires transformers >= 4.57.1. Greedy decoding and max_pixels = 262144 are already the
shipped defaults, so loading the model as-is reproduces the reported configuration — no extra
flags required.
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
"pxiaoyu/Oculora",
dtype=torch.bfloat16,
device_map="auto",
)
# max_pixels = 262144 is already the shipped default
processor = AutoProcessor.from_pretrained("pxiaoyu/Oculora")
messages = [{
"role": "user",
"content": [
{"type": "image", "url": "https://example.com/fundus.jpg"},
{"type": "text", "text": "Describe the fundus findings and give the most likely diagnosis."},
],
}]
inputs = processor.apply_chat_template(
messages,
add_generation_prompt=True,
tokenize=True,
return_dict=True,
return_tensors="pt",
).to(model.device)
# greedy decoding is already the shipped default
out = model.generate(**inputs, max_new_tokens=1000)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
vLLM — OpenAI-compatible server
python -m vllm.entrypoints.openai.api_server \
--model pxiaoyu/Oculora \
--served-model-name oculora \
--host 127.0.0.1 --port 8002 \
--dtype bfloat16 \
--max-model-len 8192 \
--gpu-memory-utilization 0.75 \
--max-num-seqs 8 \
--limit-mm-per-prompt '{"image":4}' \
--trust-remote-code
Query it as any chat-completions endpoint with image_url content parts. vLLM reads
generation_config.json from the repo, so "temperature": 0 is already the default.
Hardware — bfloat16 weights are ~17 GB plus KV cache; a single 40 GB A100 or 48 GB L40S is
comfortable at --max-model-len 8192.
Evaluation
Multiple-choice and short-answer questions assess multimodal understanding; report generation and medical consultation assess open-ended clinical generation. Every model below was evaluated under an identical protocol on the same test sets.
| Multimodal understanding | Oculora | Qwen3-VL | MedGemma | Lingshu |
|---|---|---|---|---|
| Internal MCQ · Accuracy | 0.8197 | 0.3230 | 0.3470 | 0.2993 |
| Internal MCQ · F1 | 0.8303 | 0.3393 | 0.3505 | 0.3020 |
| JSIEC · Accuracy | 0.6740 | 0.1610 | 0.1410 | 0.0650 |
| JSIEC · F1 | 0.7259 | 0.2676 | 0.2175 | 0.1651 |
| Internal short answer · Unigram F1 | 0.4869 | 0.2053 | 0.1501 | 0.2114 |
| Internal short answer · ROUGE-L | 0.4376 | 0.1758 | 0.1120 | 0.1862 |
| Internal short answer · METEOR | 0.4223 | 0.1466 | 0.1550 | 0.1695 |
| Internal short answer · BERTScore-F1 | 0.8352 | 0.7593 | 0.7116 | 0.7179 |
| OphthalVQA · Unigram F1 | 0.3189 | 0.1487 | 0.1451 | 0.0664 |
| OphthalVQA · ROUGE-L | 0.2749 | 0.1051 | 0.1016 | 0.0560 |
| OphthalVQA · METEOR | 0.2719 | 0.1818 | 0.1637 | 0.0434 |
| OphthalVQA · BERTScore-F1 | 0.7968 | 0.6945 | 0.7072 | 0.6741 |
| Open-ended clinical generation | Oculora | Qwen3-VL | MedGemma | Lingshu |
|---|---|---|---|---|
| FFA-IR · BLEU-2 | 0.1485 | 0.0917 | 0.0579 | 0.0110 |
| FFA-IR · ROUGE-L | 0.2222 | 0.1639 | 0.1040 | 0.0829 |
| FFA-IR · METEOR | 0.2393 | 0.1742 | 0.1962 | 0.0581 |
| FFA-IR · BERTScore-F1 | 0.8870 | 0.8673 | 0.8324 | 0.8492 |
| JAMA Challenge · BLEU-2 | 0.4466 | 0.2972 | 0.0703 | 0.0183 |
| JAMA Challenge · ROUGE-L | 0.5002 | 0.3106 | 0.1072 | 0.0305 |
| JAMA Challenge · METEOR | 0.5099 | 0.3613 | 0.2930 | 0.0422 |
| JAMA Challenge · BERTScore-F1 | 0.9265 | 0.8925 | 0.8563 | 0.8194 |
Best score per row in bold. Test sets: JSIEC n = 1,000 · OphthalVQA n = 600 · FFA-IR n = 766 · JAMA Challenge n = 177, plus a 22,000-sample internal test set. Baselines are open-weight vision–language models at a comparable parameter scale, run under an identical protocol on the same test sets: Qwen3-VL-8B, MedGemma-4B and Lingshu-7B. Bootstrap 95% confidence intervals for every value are reported in the paper.
Blinded clinical review
| Design | 50 cases, blinded, ranked by LLM judges and senior ophthalmologists |
| Outcome | Oculora ranked highest overall for correctness and for control of unsupported content |
| Rater agreement | 0.94 (95% CI 0.90 – 0.97) between LLM and ophthalmologist rankings |
Training
Three stages, all on the Qwen2.5-VL-7B backbone.
| Stage | Data | Configuration |
|---|---|---|
| 1 · Pretraining (PT) | 1.33 M image–text pairs | LoRA, rank 64, α = 128 |
| 2 · Supervised fine-tuning (SFT) | 360 K ophthalmic instructions | LoRA, rank 64, α = 128; vision encoder and multimodal projector frozen |
| 3 · Risk-aware policy optimization | — | Risk-aware DAPO objective in the VERL framework |
Risk-aware DAPO
| Setting | Value |
|---|---|
| Responses sampled per prompt | n = 8 |
| Global mini-batch per update | 128 prompts → 1,024 responses |
| Risk masking | top-K highest-risk generations (K = 0.1) masked globally across the whole mini-batch, before group-relative advantage computation |
| Reward — keyword | 0.5 |
| Reward — LLM-assessed | 0.4 |
| Reward — format | 0.1 |
Masking the riskiest decile globally rather than within each group is what stabilizes the gradients: the advantage baseline is never computed over generations the risk model has already flagged.
Inference settings used for all reported results
| Setting | Value |
|---|---|
| Decoding | greedy |
| Temperature | 0 |
| Max image pixels | 262,144 |
These are the shipped defaults.
preprocessor_config.jsonsetsmax_pixels = 262144andgeneration_config.jsonsetsdo_sample = false, so loading the model as-is reproduces the reported configuration — no extra flags required.
Architecture
| Class | Qwen2_5_VLForConditionalGeneration |
| Total parameters | 8.29 B |
| Language model | 28 layers · hidden 3584 · 28 heads · 4 KV heads (GQA) · SwiGLU · RMSNorm |
| Vision encoder | 32 layers · hidden 1280 → 3584 · 16 heads · patch 14 · spatial merge 2 · window 112 |
| Vocabulary | 152,064 |
| Context | 128 K positions · sliding window 32,768 |
| Position encoding | M-RoPE, mrope_section = [16, 24, 24], θ = 1,000,000 |
| Precision | bfloat16 |
| Chat template | ChatML (<|im_start|> / <|im_end|>) |
| Weights | 6 safetensors shards · 16.6 GB |
| Requires | transformers >= 4.57.1 |
Intended use and limitations
Oculora is a research artifact, intended for research on ophthalmic vision–language modelling and for retrospective evaluation on de-identified data.
- Not a medical device. Not cleared or approved by any regulatory body. It must not be used for diagnosis, triage or treatment decisions for real patients.
- Outputs can be fluent and confident while being wrong. Every clinical statement requires review by a qualified ophthalmologist.
- Performance is characterised on the benchmarks above — seven imaging modalities and 300+ conditions. Behaviour outside that distribution (other modalities, populations, languages, or image quality) is unmeasured.
- Training data inherits the demographic and disease-prevalence biases of its sources; the model may underperform on rare conditions and underrepresented groups.
- Do not send protected health information to any hosted inference endpoint.
License
Apache 2.0, inheriting the license of the Qwen2.5-VL base model. Users remain responsible for complying with the terms of the underlying base model and of every dataset used for evaluation.
Citation
The paper describing Oculora is currently under review; a full citation will be added here upon
publication. Until then, please cite this repository:
pxiaoyu/Oculora.
- Downloads last month
- 46
Model tree for pxiaoyu/Oculora
Base model
Qwen/Qwen2.5-VL-7B-Instruct