Byrne-VE - Tiny Self-Distilled Vision Encoder (39M)
Model will be ungated for open download once I am done with the base.
~39.3M ViT-style image encoder with SpikeWhale / Byrne DNA - RMSNorm, 2D-axial RoPE, QK-Norm, SwiGLU, HRM refine. First distilled from a frozen DINOv2-base teacher, then teacher-free DINO-style EMA self-distillation. Ended up closer to useful than the teacher signal alone. Drop-in vision backbone for a VLM.
image → patch-embed → 12 × (RoPE / QK-Norm / SwiGLU) blocks → HRM refine → RMSNorm
→ pooled (CLS) 512-d + 196 patch tokens (512-d each)
Architecture
| image size | 224 |
| patch size | 16 → 28×28 = 784 patches |
| hidden dim | 512 |
| depth | 12 blocks |
| heads | 8 (head_dim 64) |
| FFN | SwiGLU (8/3 mult) |
| position encoding | 2D-axial RoPE (no learned pos-embeds) |
| normalization | RMSNorm + QK-Norm |
| refinement | HRM iterative refine (2 steps, zero-init gate) |
| parameters | 39.34M (blocks 38.55M · patch-embed 0.39M · HRM 0.39M) |
Outputs (forward(images) -> dict): pooled [B,512] (global), tokens
[B,196,512] (per-patch), last_hidden_state [B,197,512].
How it was trained
Stage 1 - Distillation from DINOv2-base. Untrained student, frozen
facebook/dinov2-base teacher, ImageNet stream (timm/imagenet-1k-wds).
Loss is 1 - cosine_similarity on both global CLS and the patch-token
grid (teacher grid bicubically resized to the student's 14×14). ~50k+ steps,
AdamW, batch 32. HRM refine uses a zero-init scalar gate so it is a no-op at
start and wakes up (gate tanh climbed 0.0 → ~0.43).
End of distillation: about ~56% mean cosine to DINOv2-base (global ≈ 53.5%, patch ≈ 59.1%). Expected - student is ~2× smaller than the teacher.
Stage 2 - Teacher-free self-distillation (DINO-style). Start from the
distilled checkpoint. No external teacher. EMA copy of the encoder is the
target. Two differently-augmented views, online student matches EMA teacher
across views over learned prototypes. Collapse blocked the DINO way -
centering + sharpening, cosine-normalized frozen-early prototype layer,
soft teacher temperature. This stage moved every probe ~+5 over distilled
(peak ~step 4k; this release is that peak). Byrne-VE is the stage-2
checkpoint.
Evaluation
Frozen encoder, weighted k-NN, no fine-tune, same preprocess for every
model. k=20.
vs. the DINOv2-base teacher (full-bank CIFAR-10, 50k bank / 10k query)
| Model | Params | Top-1 |
|---|---|---|
| DINOv2-base (teacher) | 86M | 98.47% |
| Byrne-VE (this model) | 39M | 86.38% |
| distilled-only (stage 1) | 39M | 81.39% |
→ ~88% of the teacher's accuracy at ~45% of the params. Self-distillation added +4.99 over distilled-only.
Self-distillation gain across probes
| Probe | Byrne-VE | distilled-only | Δ |
|---|---|---|---|
| CIFAR-10, full-bank (50k bank, 10-way) | 86.38% | 81.39% | +4.99 |
| CIFAR-10, small (5k bank, 10-way) | 81.60% | 75.40% | +6.20 |
| CIFAR-100 (10k bank, 100-way) | 54.35% | 49.50% | +4.85 |
Consistent ~+5 on easy, full-bank, and 100-way. Not a probe-specific artifact.
Health checks: HRM gate tanh ≈ 0.355 (alive, off-zero) · no NaN/Inf in any
weight · determinism cos 1.0000 · cross-image pooled cos 0.142
(discriminative).
Usage
pip install -r requirements.txt
python -m vision.infer --ckpt byrne_ve.pt --image photo.jpg
python -m vision.eval_probe --ckpt byrne_ve.pt # reproduce CIFAR-10 k-NN
from vision.infer import load_encoder
from vision.dataset import load_image
m, cfg = load_encoder("byrne_ve.pt")
x = load_image("photo.jpg", cfg.image_size).unsqueeze(0)
out = m(x)
emb = out["pooled"] # [1, 512] global image embedding
patches = out["tokens"] # [1, 196, 512] per-patch features
Citation
@misc{byrne_ve_2026,
title = {Byrne-VE: A Tiny Self-Distilled Vision Encoder},
author = {Quazim0t0},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/Quazim0t0/Byrne-VE}},
note = {39M ViT-style encoder distilled from DINOv2-base, then improved
teacher-free via DINO-style self-distillation.}
}
License
Apache-2.0.
Escarda vs Byrne - vision family comparison
Byrne = HRM refine. Escarda = Byrne + JEPA on the vision encoder and the LM trunk. Auxiliary only. Zero inference cost.
Vision encoder (DINOv2 teacher-alignment, n=1024 held-out):
| Byrne-VE | Escarda-VE | |
|---|---|---|
| Params | 39.34M | 39.60M (+JEPA head) |
| CLS cosine | 0.776 | 0.771 |
| PATCH cosine | 0.600 | 0.584 |
| JEPA self-consistency | - | 0.040 |
Docling (same held-out doc images, atomic DocTags): both emit well-formed
DocTags. Byrne-Docling is a bit more complete on the hardest samples (closes
</formula>, includes the <code> wrapper), which matches the slightly higher
teacher-alignment. Escarda-Docling is structurally on par and has the JEPA
representation-learning trait.
Pros/cons. Byrne (HRM): higher teacher-alignment, all capacity on distillation fidelity; no self-supervised objective. Escarda (HRM+JEPA): self-supervised neighbour-prediction (richer spatial structure) at zero inference cost, trading ~1-3% teacher-alignment. Same size class.
Family repos: Byrne-VE · Escarda-VE · Byrne-Docling-131M · Escarda-Docling-126M
v2 update (448px)
This encoder is now the 448px variant (patch16, 28×28 = 784 tokens) - higher input resolution for finer detail (text/documents). Same 39.3M architecture as v1; only input resolution and token count changed (params are resolution-independent). Distilled from DINOv2-base at 448px then DINO-style self-distilled. This is the encoder Byrne-VLM-131M uses now. (v1 was 224px / 196 tokens.)