Cosmos3-Nano-Policy-DROID, FP8 + RTN + moving-average activation observer (ModelOpt W8A8, per-tensor static)

FP8 checkpoint of nvidia/Cosmos3-Nano-Policy-DROID (revision 805c0d6d46196ecdc789213a6de72303c383895b) in NVIDIA's Cosmos3-Nano@fp8 layout, one of the three builds selected from an np-quantizer-v2 algorithm sweep. Siblings: np-cr-test/Cosmos3-Nano-Policy-DROID-FP8 (plain RTN, the NVIDIA-identical build) and the other selected build. The weights are the delivered RTN build's; only the static input scales change: each is a moving average of the per-call maxima rather than their maximum, so it clips the rarest activations.

Scheme

Algorithm none (RTN weights); activation scales from MOVING_AVERAGE_MINMAX (averaging constant 0.01) instead of the global max
Quantized 504 Linear layers: q/k/v/o and gate/up/down of both towers in all 36 decoder layers
Weights FP8 E4M3, one scale per tensor
Activations FP8 E4M3, one static scale per tensor, calibrated on 256 transformer calls captured from NVIDIA's policy server
Kept in BF16 proj_in, proj_out, time_embedder, action_proj_in, action_proj_out, lm_head, the vision encoder, embeddings and norms
Format ModelOpt 0.44 FP8 in NVIDIA's layout (same quantization_config, quantizer buffers and modelopt_state.pth as nvidia/Cosmos3-Nano@fp8)

How it was selected

Selected on 2026-09-23 from the first seven builds of the same scheme, scored by a fake-quant replay against the BF16 original on 32 DROID observations: the campaign gate (action error overall <= 0.13 sigma, gripper-state changes <= 0.5% of steps), then the lowest overall. The comparison was later extended to 22 builds, five of them served on NVIDIA's policy server; all tables are in the section "알고리즘 비교 결과 (2026-09-24)" below.

Verification

Check Result
Files, tensor names, dtypes and shapes vs nvidia/Cosmos3-Nano@fp8 3329 common tensors, 0 dtype or shape differences; the rest is NVIDIA's docs, the base model's audio branch and 5 np-quantizer-v2 sidecars
quantization_config, root hf_quant_config.json Identical to NVIDIA's
modelopt_state.pth, transformer and root Same configurations; 1527/509/504 and 1478/253/252 common entries, 0 differ
Values Differ from NVIDIA's pipeline by design (the input scales come from a different observer); FP8 codes follow ModelOpt's compress rounding as in the RTN build

How far it is from the BF16 original

Comparison Path overall (sigma) [95% CI] x eager-vs-compile joint MAE (deg) SNR (dB) gripper agree [95% CI]
This build (RTN + moving-average observer), FP8 vs BF16 original NVIDIA policy server, FP8 GEMM, same seeds 0.099 [0.083, 0.117] 0.76 1.96 19.5 99.32% [98.60, 99.67]
This build, fake quant vs BF16 original fakequant_loader, offline replay 0.105 [0.080, 0.131] 0.81 1.94 19.0 99.61% [99.00, 99.85]
RTN build, FP8 vs BF16 original NVIDIA policy server, FP8 GEMM, same seeds 0.127 [0.095, 0.156] 0.98 2.28 17.4 98.83% [97.96, 99.33]
BF16 eager vs BF16 torch.compile (yardstick) policy server, same seeds 0.130 [0.079, 0.175] 1.00 2.37 15.2 100.00% [99.26, 100.00]
BF16 seed vs BF16 other seed (yardstick) policy server, 4 obs x 4 seeds 0.392 [0.306, 0.474] 3.02 8.04 5.5 100.00% [99.50, 100.00]

알고리즘 비교 결과 (2026-09-24)

이 저장소는 아래 표의 RTN + 이동평균 observer 0.01 (전달) 행입니다. (전달)이 붙은 세 build가 HF에 올린 저장소입니다.

표의 build 저장소
RTN (전달) np-cr-test/Cosmos3-Nano-Policy-DROID-FP8
SQ auto-α (전달) np-cr-test/Cosmos3-Nano-Policy-DROID-FP8-SmoothQuant
RTN + 이동평균 observer 0.01 (전달) np-cr-test/Cosmos3-Nano-Policy-DROID-FP8-RTN-MovingAvgObserver

결론

  • NVIDIA 형식(W8A8 FP8, per-tensor static)에서는 RTN보다 확실히 나은 알고리즘이나 조합이 없습니다.
  • fake quant와 실제 서빙에서 순위가 서로 반대입니다. 서빙에서 RTN과 벌어진 차이도 대부분 관측 2개에서 나옵니다. 그래서 관측 32개로는 build 간 순위를 가를 수 없습니다.
  • 전달본 3개(RTN, SQ auto-α, 이동평균 observer)는 그대로 둡니다.
  • per-channel weight / per-token activation 계열은 fake quant 오차가 가장 낮습니다. 그러나 RTN과 유의한 차이가 없고, NVIDIA 서버가 읽는 형식도 아닙니다.

1. 실험 설정

항목 값
양자화 대상 transformer의 Linear 504개 (두 tower × 36 layer × q/k/v/o, gate/up/down). 스킵 목록은 위 Scheme 표
scheme FP8 E4M3 W8A8, weight per-tensor, activation per-tensor static (nvidia/Cosmos3-Nano@fp8과 같음)
export MODELOPT_FP8, NVIDIA diffusers 레이아웃
알고리즘 데이터 관측당 transformer 호출 1개씩, 32개
activation calibration NVIDIA 정책 서버 호출 256개 (관측 32 × denoising 4 step × CFG 2)
알고리즘 파라미터 AQ 기본값. 바꾼 값은 build 이름에 적음 (SQ = SmoothQuant)
평가 관측 DROID 32개. calibration과 같은 관측이라 in-sample
오차 지표 overall: 정규화 action RMS (DROID action σ 단위). 판정 변화: 32 관측 × 32 step 중 gripper open/closed 판정이 BF16과 다른 step 수
통계 관측 단위 bootstrap 10,000회. RTN과의 차이는 같은 resample로 짝지어 계산
gate overall ≤ 0.13σ, 판정 변화 ≤ 0.5% (≤ 5/1024)
참고 크기 BF16 eager vs torch.compile 0.130σ, BF16 시드만 바꿀 때 0.392σ
평가 경로 실행 기준 대상
fake quant FP8 코드를 dequantize한 가중치 + 입력 FP8 quantize-dequantize (AQ FAKE_QUANT 평가와 비트 동일) FP adapter의 BF16 출력 22개 전부
서빙 NVIDIA cosmos-framework 정책 서버, TorchAO FP8 커널 (2 × RTX 4090, FSDP2 + CFG 병렬) 같은 서버의 BF16 출력 5개

2. 서빙 결과

전달본 3개와, 전달 후 추가한 NVIDIA 형식 build 8개 중 fake quant 오차가 가장 낮은 SQ auto-α + GPTQ, 판정 변화가 가장 적은 SQ auto-α + AutoRound를 서빙했습니다. 다섯 build 모두 서버가 FP8 가중치 504/504개를 TorchAO로 로드했습니다.

build 서빙 σ [95% CI] 서빙 RTN 대비 [95% CI] 서빙 판정 변화 fake quant σ fake quant 판정 변화 gate (fake quant / 서빙)
RTN (전달) 0.127 [0.095, 0.156] - 12 0.097 9 초과 / 초과
RTN + 이동평균 observer 0.01 (전달) 0.099 [0.083, 0.117] -0.028 [-0.051, -0.004] 7 0.105 4 통과 / 초과
SQ auto-α + GPTQ 0.104 [0.082, 0.127] -0.023 [-0.049, +0.003] 8 0.098 8 초과 / 초과
SQ auto-α (전달) 0.106 [0.079, 0.137] -0.021 [-0.047, +0.005] 4 0.100 4 통과 / 통과
SQ auto-α + AutoRound 0.116 [0.084, 0.149] -0.011 [-0.034, +0.010] 5 0.109 2 통과 / 통과

차이가 어디서 나오는지 관측별로 나눠 봤습니다.

build 서빙 RTN 대비 RTN 대비 차이를 가장 키운 관측 2개 그 2개를 뺀 RTN 대비
RTN + 이동평균 observer 0.01 (전달) -0.028 28, 2 -0.012
SQ auto-α (전달) -0.021 5, 28 -0.006
SQ auto-α + GPTQ -0.023 5, 28 -0.005
SQ auto-α + AutoRound -0.011 28, 8 -0.000
  • 네 build 모두 관측 28번이 차이를 키운 관측 2개에 들어 있습니다.
  • 서빙 RTN은 오차 상위 관측 3개(5, 28, 2)가 제곱오차의 46%를 차지합니다. 같은 체크포인트를 fake quant로 돌리면 상위 관측이 7, 12, 15번으로 바뀝니다.
  • 이동평균 observer는 95% CI가 0을 배제합니다. 다만 4개를 동시에 비교한 것을 Bonferroni로 보정하면 p ≈ 0.09입니다.

3. fake quant 결과

3.1 NVIDIA 형식 16개

순위 build σ [95% CI] RTN 대비 [95% CI] 판정 변화 /1024
1 RTN (전달) 0.097 [0.079, 0.116] - 9
2 SQ auto-α + GPTQ 0.098 [0.078, 0.120] +0.001 [-0.022, +0.025] 8
3 SQ auto-α (전달) 0.100 [0.078, 0.122] +0.002 [-0.023, +0.028] 4
4 SQ auto-α, norm-FC 쌍만 0.103 [0.078, 0.131] +0.006 [-0.021, +0.035] 8
5 SQ α=0.5 0.103 [0.075, 0.136] +0.006 [-0.024, +0.040] 12
6 RTN + 이동평균 observer 0.05 0.104 [0.082, 0.127] +0.007 [-0.014, +0.031] 9
7 RTN + 이동평균 observer 0.01 (전달) 0.105 [0.080, 0.131] +0.008 [-0.016, +0.033] 4
8 GPTQ 0.106 [0.078, 0.133] +0.009 [-0.020, +0.037] 7
9 SQ α=0.3 0.108 [0.084, 0.132] +0.011 [-0.009, +0.033] 7
10 SQ auto-α + AutoRound 0.109 [0.082, 0.136] +0.012 [-0.010, +0.036] 2
11 AWQ + GPTQ 0.111 [0.085, 0.136] +0.014 [-0.006, +0.037] 12
12 SQ auto-α + 이동평균 observer 0.01 0.119 [0.087, 0.152] +0.022 [-0.012, +0.057] 7
13 AutoRound 0.121 [0.083, 0.160] +0.024 [-0.015, +0.065] 8
14 RTN + RANGE_SHRINK 0.132 [0.095, 0.166] +0.035 [+0.001, +0.068] 21
15 AWQ 0.135 [0.092, 0.175] +0.038 [-0.003, +0.077] 9
16 RTN + HISTOGRAM observer 0.212 [0.185, 0.240] +0.115 [+0.085, +0.143] 17
  • NVIDIA 형식 중 RTN보다 오차가 낮은 build는 없습니다.
  • AutoRound는 AQ 구현상 INT, UINT, NVFP4 가중치만 튜닝합니다. 그래서 "AutoRound"와 "SQ auto-α + AutoRound" build의 FP8 코드는 각각 RTN, SQ auto-α와 바이트 단위로 같습니다(139억 개 중 0개 차이). 두 행이 RTN, SQ auto-α와 다른 것은 AutoRound 단계 뒤 activation calibration이 달라졌기 때문입니다(입력 scale 504개 중 250개, 302개). 즉 이 두 행은 AutoRound의 효과가 아닙니다.
  • HISTOGRAM observer는 확실히 나쁩니다. RANGE_SHRINK는 보정 전 기준으로만 나쁩니다(21개 비교 보정 시 유의하지 않음).
  • GPTQ는 AQ schema가 per-tensor weight를 막습니다. 실험 브랜치에서 이 검사를 풀었고, 원소 1개짜리 scale을 행 수만큼 펴는 수정도 넣었습니다. 이 수정이 없으면 Scale shape torch.Size([1, 1]) is incompatible with input shape torch.Size([4096, 1])로 실패합니다.

3.2 per-channel weight / per-token activation 6개 (탐색용)

RTN 대비 차이는 NVIDIA 형식 RTN(per-tensor)과 비교한 값입니다.

순위 build σ [95% CI] RTN 대비 [95% CI] 판정 변화 /1024
1 SQ auto-α + GPTQ 0.088 [0.066, 0.115] -0.009 [-0.037, +0.021] 7
2 RTN 0.090 [0.069, 0.112] -0.008 [-0.019, +0.003] 8
3 AutoRound 0.091 [0.073, 0.108] -0.006 [-0.026, +0.012] 8
4 GPTQ 0.095 [0.067, 0.128] -0.002 [-0.032, +0.033] 4
5 AWQ 0.096 [0.076, 0.114] -0.002 [-0.014, +0.011] 5
6 SQ auto-α 0.103 [0.078, 0.128] +0.006 [-0.014, +0.027] 17

이 계열은 두 가지 이유로 지금 형식으로는 전달할 수 없습니다.

  • NVIDIA 서버의 ModelOpt 체크포인트 로더(apply_modelopt_fp8_checkpoint_inplace)는 두 scale을 (1, 1)로 읽어 PerTensor()로만 설치합니다.
  • 같은 scheme은 프레임워크 런타임 FP8(quantization.method=fp8, fp8_granularity=per_row)에만 있습니다. 이 경로는 BF16 가중치를 로드할 때 RTN으로 양자화하므로, AQ 알고리즘 결과를 담을 수 없습니다.

4. 결정과 남은 일

항목 내용
전달본 RTN, SQ auto-α, 이동평균 observer 유지. 두 경로 모두 gate를 통과한 전달본은 SQ auto-α뿐이고, RTN은 NVIDIA와 비트 동일한 기준으로 둠
SQ auto-α + AutoRound 두 경로 모두 gate 통과. 하지만 SQ auto-α보다 오차가 낮지 않아 추가하지 않음. 가중치는 SQ auto-α와 같음(AutoRound는 FP8을 튜닝하지 않음)
SQ auto-α + GPTQ 두 경로 모두 판정 변화 8로 gate 초과
순위 확정에 필요한 것 calibration과 lab이 겹치지 않는 관측으로 다시 평가. /workspace/cosmos3-policy-quant/splits/splits.json(내부 서버)의 selection_validation(AUTOLab, IPRL, PennPAL, RPL)을 쓰면 됨. 지금 32개는 전부 calibration lab(TRI, CLVR, IRIS, ILIAD, RAD) 소속

5. 제약

  • 평가가 in-sample입니다. calibration, 알고리즘 데이터, 선정이 모두 같은 32개 관측에서 나왔습니다.
  • AQ는 DIFFUSION 모델에 SVDQuant만 허용합니다. 비교는 이 게이트를 푼 실험 브랜치 cosmos3-algo-sweep(미푸시)에서 했습니다. 커밋은 a3a1b56c4(게이트), ea30283f3(GPTQ per-tensor 검사), 8cdadc414(scale shape)입니다.
  • vLLM-Omni 서빙은 해보지 않았습니다.

6. 재현 (내부 GPU 서버 경로)

항목 위치 (/workspace/cosmos3-policy-quant/)
config fq/configs/algo/*.yaml, fq/configs/algo3/*.yaml
실행 스크립트 fq/algo_sweep*_run.sh, fq/algo_served4_*_run.sh
순위표 fq/rank_all.sh → fq/eval/sweep_all.md, fq/rank_served.sh → fq/eval/served_all.md
평가 하네스 harness/fakequant_loader.py, eval_fakequant.py, server_fidelity.py, stats_table.py, sweep_table.py

Use

Served exactly as the RTN build: cosmos-framework's RoboLab action policy server with checkpoint_path set to this repository and mixed_precision_first_steps=0, mixed_precision_last_steps=0, mixed_precision_w8a16_cache="none" (generation under torch.no_grad, see the RTN build's card). fakequant_loader.py evaluates it as a fake-quantized model on any GPU.

Caveats

  • In-sample twice over: the calibration calls and the selection both use the 32 scored observations, so the selected builds' margin over the others is optimistic.
  • vLLM-Omni was not run; its step policy would serve W8A16 on the policy's 4 denoising steps.

Provenance

Source nvidia/Cosmos3-Nano-Policy-DROID@805c0d6d46196ecdc789213a6de72303c383895b (BF16)
Quantizer np-quantizer-v2 experiment branch cosmos3-algo-sweep commit a3a1b56c4 (integration branch 411651222 + the DIFFUSION algorithm gate lifted), MODELOPT_FP8 in the diffusers layout; config in verification/config.yaml
ModelOpt state files nvidia-modelopt 0.44.0

License

A quantized derivative of nvidia/Cosmos3-Nano-Policy-DROID, distributed under the same OpenMDW License 1.1. See NVIDIA's model card for intended use and limitations.

Downloads last month
14
Safetensors
Model size
16B params
Tensor type
BF16
·
F8_E4M3
·
Video Preview
loading

Model tree for np-cr/Cosmos3-Nano-Policy-DROID-FP8-RTN-MovingAvgObserver

Quantized
(11)
this model