Twin Loss Curves β€” and the Seed-Null That Retracted the Headline

Correction (13 July 2026). The first version of this card reported that the two models below compute measurably differently inside β€” a "4.4Γ— quieter internals under int8" figure, an "84% of depth" divergence, and a routing-stability edge. We then ran the one control we had flagged as "obvious next": an independent-seed null. We trained three more models and re-scanned. Every one of those differential-internal claims fell inside the noise between two ordinary runs, and is retracted. This card now documents the original claims, the control that refuted them, and what survived. The retraction is the point β€” see the full post-mortem at https://www.tetracta.ai/note-xray-twin-study.html.

Update (17 July 2026). We re-ran both arms through the current production X-Ray pipeline (rs-1.3, unmodified β€” the same scanner outside teams now use). The character of the observation is unchanged: under int8 both twins keep byte-identical outputs (0/6 behavior change on either), while their per-layer internal responses resolve differently and deterministically. Per the seed-null control below, we attribute none of this to the architecture β€” what it demonstrates is the instrument: per-pair internal resolution that output-level evals cannot provide by construction. The scanner is now open to outside teams (20 free trial seats): https://www.tetracta.ai/xray.html

Two 0.93B language models, trained under strictly matched conditions β€” identical seed, identical data in identical order, identical schedule β€” differing in exactly one component: the attention weighting function (standard softmax vs Tetracta's undisclosed rational operator, zero extra params).

vanilla/ rational/
val bits-per-byte @ step 30 000 1.0092 1.0102
parameters 0.929 B 0.929 B

By every output-level metric these two are twins β€” and the seed-null control made that the strongest claim here, not the weakest (below).

What we first claimed, and what the control did to it

We scanned both with Tetracta Model X-Ray and published four internal findings. A reviewer asked for the denominator: how much do two runs of the same recipe, differing only in seed, already diverge inside? We trained VAN-seed-B, VAN-seed-C (two more softmax runs) and RAT-seed-B, and measured, with thresholds sealed in advance. Five checkpoints; ~$300 compute; all pods terminated; all checkpoints md5-sealed.

Loss twins β€” CONFIRMED, and strengthened. The three softmax seeds landed at bpb 1.0092 / 1.0072 / 1.0085; the two rational seeds at 1.0102 / 1.0101. The spread between softmax seeds (0.0020 bpb) is larger than the original softmax-vs-rational gap (0.0010). The operator moves loss by less than the seed does β€” a leaderboard genuinely cannot separate them.

"4.4Γ— quieter inside" β€” RETRACTED. Internal disturbance under int8, per model: VAN-42 4.166, VAN-B 0.573, VAN-C 0.967, RAT-42 1.277, RAT-B 1.104. Spread among softmax seeds alone: 7.3Γ—. The original VAN-42 was a high outlier; the softmax median (0.97) sits below rational (1.19). The direction did not survive, let alone the 4.4Γ—.

Routing-stability edge β€” RETRACTED. int8 token-flip rate: VAN-42 2.97%, VAN-B 0.93%, VAN-C 1.74%, RAT-42 0.52%, RAT-B 2.66%. The rational replication lands inside the softmax band; RAT-42's low value was a lucky draw.

"84% of depth" divergence β€” RETRACTED (as a ceiling). Cross-scan mean deviation: operator pair 0.000307, seed pair 0.000324 β€” both span 84% of depth (ratio 1.06Γ—). Two ordinary softmax runs diverge internally exactly as much as the operator swap does. Our own note had flagged 84% as a ceiling indicator; the control confirmed it, at our expense.

A null we kept (unchanged). An early read suggested rational abstains more on trick questions. Adversarial re-verification killed it on day one (a generation-budget artifact, pβ‰ˆ0.2). Still a null.

What survived, and is seed-independent

On a converged public MoE (Qwen1.5-MoE-A2.7B, 1.47M routing decisions), real int8 kernels flip about twice the routing decisions of the round-to-nearest simulation most quantization audits use (~2.6% vs ~1.3%, deterministic). Anyone auditing quant with a simulation sees roughly half the real kernel's routing effect. This does not depend on the twins, and it stands.

And the capability the retraction does not touch: the instrument scanned a bespoke, undisclosed architecture unchanged β€” the product pipeline ran on a custom operator, not a HuggingFace model.

Honest limits

0.93B models at step 30k of 157k β€” young. Simulated per-row RTN quantization, not production GPTQ/AWQ. Behavioural comparisons rest on n=6 probe prompts. Nothing here is a scaling law, and nothing here claims rational produces a better model. The internal-difference claims we did make were retracted by our own control β€” which is the methodology this release is really about.

Files

vanilla/  model-step{10000,20000,30000}.safetensors + config.json   # softmax baseline
rational/ model-step{10000,20000,30000}.safetensors + config.json   # Tetracta rational attention
xray/     portrait-{vanilla,rational}-step30000.png
xray_summary.json    weight_error.json + reproduce_weight_error.py    manifest.json (sha256)

float32. lm_head is tied to tok_emb. rational/ cannot be run correctly without the undisclosed operator (method patent application in preparation) β€” loading it into a softmax model yields meaningless output; please do not benchmark that as "Tetracta rational."

The point of all this

Output-level equality is not evidence of internal equality β€” but a single internal measurement is not evidence of a real effect either. The discipline that separates a real signal from the luck of a seed is a pre-registered negative control, run even when it overturns your own headline. That discipline, not any one number, is what Tetracta sells.

Tetracta AI Teams β€” for humans, like humans.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support