Byrne-VE - Tiny Self-Distilled Vision Encoder (39M)

Model will be ungated for open download once I am done with the base.

~39.3M ViT-style image encoder with SpikeWhale / Byrne DNA - RMSNorm, 2D-axial RoPE, QK-Norm, SwiGLU, HRM refine. First distilled from a frozen DINOv2-base teacher, then teacher-free DINO-style EMA self-distillation. Ended up closer to useful than the teacher signal alone. Drop-in vision backbone for a VLM.

image → patch-embed → 12 × (RoPE / QK-Norm / SwiGLU) blocks → HRM refine → RMSNorm
      → pooled (CLS) 512-d   +   196 patch tokens (512-d each)

Architecture

image size 224
patch size 16 → 28×28 = 784 patches
hidden dim 512
depth 12 blocks
heads 8 (head_dim 64)
FFN SwiGLU (8/3 mult)
position encoding 2D-axial RoPE (no learned pos-embeds)
normalization RMSNorm + QK-Norm
refinement HRM iterative refine (2 steps, zero-init gate)
parameters 39.34M (blocks 38.55M · patch-embed 0.39M · HRM 0.39M)

Outputs (forward(images) -> dict): pooled [B,512] (global), tokens [B,196,512] (per-patch), last_hidden_state [B,197,512].

How it was trained

Stage 1 - Distillation from DINOv2-base. Untrained student, frozen facebook/dinov2-base teacher, ImageNet stream (timm/imagenet-1k-wds). Loss is 1 - cosine_similarity on both global CLS and the patch-token grid (teacher grid bicubically resized to the student's 14×14). ~50k+ steps, AdamW, batch 32. HRM refine uses a zero-init scalar gate so it is a no-op at start and wakes up (gate tanh climbed 0.0 → ~0.43).

End of distillation: about ~56% mean cosine to DINOv2-base (global ≈ 53.5%, patch ≈ 59.1%). Expected - student is ~2× smaller than the teacher.

Stage 2 - Teacher-free self-distillation (DINO-style). Start from the distilled checkpoint. No external teacher. EMA copy of the encoder is the target. Two differently-augmented views, online student matches EMA teacher across views over learned prototypes. Collapse blocked the DINO way - centering + sharpening, cosine-normalized frozen-early prototype layer, soft teacher temperature. This stage moved every probe ~+5 over distilled (peak ~step 4k; this release is that peak). Byrne-VE is the stage-2 checkpoint.

Evaluation

Frozen encoder, weighted k-NN, no fine-tune, same preprocess for every model. k=20.

vs. the DINOv2-base teacher (full-bank CIFAR-10, 50k bank / 10k query)

Model Params Top-1
DINOv2-base (teacher) 86M 98.47%
Byrne-VE (this model) 39M 86.38%
distilled-only (stage 1) 39M 81.39%

→ ~88% of the teacher's accuracy at ~45% of the params. Self-distillation added +4.99 over distilled-only.

Self-distillation gain across probes

Probe Byrne-VE distilled-only Δ
CIFAR-10, full-bank (50k bank, 10-way) 86.38% 81.39% +4.99
CIFAR-10, small (5k bank, 10-way) 81.60% 75.40% +6.20
CIFAR-100 (10k bank, 100-way) 54.35% 49.50% +4.85

Consistent ~+5 on easy, full-bank, and 100-way. Not a probe-specific artifact.

Health checks: HRM gate tanh ≈ 0.355 (alive, off-zero) · no NaN/Inf in any weight · determinism cos 1.0000 · cross-image pooled cos 0.142 (discriminative).

Usage

pip install -r requirements.txt
python -m vision.infer --ckpt byrne_ve.pt --image photo.jpg
python -m vision.eval_probe --ckpt byrne_ve.pt    # reproduce CIFAR-10 k-NN
from vision.infer import load_encoder
from vision.dataset import load_image
m, cfg = load_encoder("byrne_ve.pt")
x = load_image("photo.jpg", cfg.image_size).unsqueeze(0)
out = m(x)
emb     = out["pooled"]   # [1, 512]   global image embedding
patches = out["tokens"]   # [1, 196, 512] per-patch features

Citation

@misc{byrne_ve_2026,
  title        = {Byrne-VE: A Tiny Self-Distilled Vision Encoder},
  author       = {Quazim0t0},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Quazim0t0/Byrne-VE}},
  note         = {39M ViT-style encoder distilled from DINOv2-base, then improved
                  teacher-free via DINO-style self-distillation.}
}

License

Apache-2.0.

Escarda vs Byrne - vision family comparison

Byrne = HRM refine. Escarda = Byrne + JEPA on the vision encoder and the LM trunk. Auxiliary only. Zero inference cost.

Vision encoder (DINOv2 teacher-alignment, n=1024 held-out):

Byrne-VE Escarda-VE
Params 39.34M 39.60M (+JEPA head)
CLS cosine 0.776 0.771
PATCH cosine 0.600 0.584
JEPA self-consistency - 0.040

Docling (same held-out doc images, atomic DocTags): both emit well-formed DocTags. Byrne-Docling is a bit more complete on the hardest samples (closes </formula>, includes the <code> wrapper), which matches the slightly higher teacher-alignment. Escarda-Docling is structurally on par and has the JEPA representation-learning trait.

Pros/cons. Byrne (HRM): higher teacher-alignment, all capacity on distillation fidelity; no self-supervised objective. Escarda (HRM+JEPA): self-supervised neighbour-prediction (richer spatial structure) at zero inference cost, trading ~1-3% teacher-alignment. Same size class.

Family repos: Byrne-VE · Escarda-VE · Byrne-Docling-131M · Escarda-Docling-126M

v2 update (448px)

This encoder is now the 448px variant (patch16, 28×28 = 784 tokens) - higher input resolution for finer detail (text/documents). Same 39.3M architecture as v1; only input resolution and token count changed (params are resolution-independent). Distilled from DINOv2-base at 448px then DINO-style self-distilled. This is the encoder Byrne-VLM-131M uses now. (v1 was 224px / 196 tokens.)

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Quazim0t0/Byrne-VE