Cosmos3-Nano-Policy-DROID, FP8 + RTN + moving-average activation observer (ModelOpt W8A8, per-tensor static)
FP8 checkpoint of nvidia/Cosmos3-Nano-Policy-DROID
(revision 805c0d6d46196ecdc789213a6de72303c383895b) in NVIDIA's Cosmos3-Nano@fp8 layout, one of the three builds
selected from an np-quantizer-v2 algorithm sweep. Siblings: np-cr-test/Cosmos3-Nano-Policy-DROID-FP8 (plain RTN, the
NVIDIA-identical build) and the other selected build. The weights are the delivered RTN build's; only the static input scales change: each is a moving average of the per-call maxima rather than their maximum, so it clips the rarest activations.
Scheme
|
|
| Algorithm |
none (RTN weights); activation scales from MOVING_AVERAGE_MINMAX (averaging constant 0.01) instead of the global max |
| Quantized |
504 Linear layers: q/k/v/o and gate/up/down of both towers in all 36 decoder layers |
| Weights |
FP8 E4M3, one scale per tensor |
| Activations |
FP8 E4M3, one static scale per tensor, calibrated on 256 transformer calls captured from NVIDIA's policy server |
| Kept in BF16 |
proj_in, proj_out, time_embedder, action_proj_in, action_proj_out, lm_head, the vision encoder, embeddings and norms |
| Format |
ModelOpt 0.44 FP8 in NVIDIA's layout (same quantization_config, quantizer buffers and modelopt_state.pth as nvidia/Cosmos3-Nano@fp8) |
How it was selected
Selected on 2026-09-23 from the first seven builds of the same scheme, scored by a fake-quant replay against the
BF16 original on 32 DROID observations: the campaign gate (action error overall <= 0.13 sigma, gripper-state changes
<= 0.5% of steps), then the lowest overall. The comparison was later extended to 22 builds, five of them served on
NVIDIA's policy server; all tables are in the section "알고리즘 비교 결과 (2026-09-24)" below.
Verification
| Check |
Result |
Files, tensor names, dtypes and shapes vs nvidia/Cosmos3-Nano@fp8 |
3329 common tensors, 0 dtype or shape differences; the rest is NVIDIA's docs, the base model's audio branch and 5 np-quantizer-v2 sidecars |
quantization_config, root hf_quant_config.json |
Identical to NVIDIA's |
modelopt_state.pth, transformer and root |
Same configurations; 1527/509/504 and 1478/253/252 common entries, 0 differ |
| Values |
Differ from NVIDIA's pipeline by design (the input scales come from a different observer); FP8 codes follow ModelOpt's compress rounding as in the RTN build |
How far it is from the BF16 original
| Comparison |
Path |
overall (sigma) [95% CI] |
x eager-vs-compile |
joint MAE (deg) |
SNR (dB) |
gripper agree [95% CI] |
| This build (RTN + moving-average observer), FP8 vs BF16 original |
NVIDIA policy server, FP8 GEMM, same seeds |
0.099 [0.083, 0.117] |
0.76 |
1.96 |
19.5 |
99.32% [98.60, 99.67] |
| This build, fake quant vs BF16 original |
fakequant_loader, offline replay |
0.105 [0.080, 0.131] |
0.81 |
1.94 |
19.0 |
99.61% [99.00, 99.85] |
| RTN build, FP8 vs BF16 original |
NVIDIA policy server, FP8 GEMM, same seeds |
0.127 [0.095, 0.156] |
0.98 |
2.28 |
17.4 |
98.83% [97.96, 99.33] |
| BF16 eager vs BF16 torch.compile (yardstick) |
policy server, same seeds |
0.130 [0.079, 0.175] |
1.00 |
2.37 |
15.2 |
100.00% [99.26, 100.00] |
| BF16 seed vs BF16 other seed (yardstick) |
policy server, 4 obs x 4 seeds |
0.392 [0.306, 0.474] |
3.02 |
8.04 |
5.5 |
100.00% [99.50, 100.00] |
알고리즘 비교 결과 (2026-09-24)
이 저장소는 아래 표의 RTN + 이동평균 observer 0.01 (전달) 행입니다. (전달)이 붙은 세 build가 HF에 올린 저장소입니다.
| 표의 build |
저장소 |
| RTN (전달) |
np-cr-test/Cosmos3-Nano-Policy-DROID-FP8 |
| SQ auto-α (전달) |
np-cr-test/Cosmos3-Nano-Policy-DROID-FP8-SmoothQuant |
| RTN + 이동평균 observer 0.01 (전달) |
np-cr-test/Cosmos3-Nano-Policy-DROID-FP8-RTN-MovingAvgObserver |
결론
- NVIDIA 형식(W8A8 FP8, per-tensor static)에서는 RTN보다 확실히 나은 알고리즘이나 조합이 없습니다.
- fake quant와 실제 서빙에서 순위가 서로 반대입니다. 서빙에서 RTN과 벌어진 차이도 대부분 관측 2개에서 나옵니다. 그래서 관측 32개로는 build 간 순위를 가를 수 없습니다.
- 전달본 3개(RTN, SQ auto-α, 이동평균 observer)는 그대로 둡니다.
- per-channel weight / per-token activation 계열은 fake quant 오차가 가장 낮습니다. 그러나 RTN과 유의한 차이가 없고, NVIDIA 서버가 읽는 형식도 아닙니다.
1. 실험 설정
| 항목 |
값 |
| 양자화 대상 |
transformer의 Linear 504개 (두 tower × 36 layer × q/k/v/o, gate/up/down). 스킵 목록은 위 Scheme 표 |
| scheme |
FP8 E4M3 W8A8, weight per-tensor, activation per-tensor static (nvidia/Cosmos3-Nano@fp8과 같음) |
| export |
MODELOPT_FP8, NVIDIA diffusers 레이아웃 |
| 알고리즘 데이터 |
관측당 transformer 호출 1개씩, 32개 |
| activation calibration |
NVIDIA 정책 서버 호출 256개 (관측 32 × denoising 4 step × CFG 2) |
| 알고리즘 파라미터 |
AQ 기본값. 바꾼 값은 build 이름에 적음 (SQ = SmoothQuant) |
| 평가 관측 |
DROID 32개. calibration과 같은 관측이라 in-sample |
| 오차 지표 |
overall: 정규화 action RMS (DROID action σ 단위). 판정 변화: 32 관측 × 32 step 중 gripper open/closed 판정이 BF16과 다른 step 수 |
| 통계 |
관측 단위 bootstrap 10,000회. RTN과의 차이는 같은 resample로 짝지어 계산 |
| gate |
overall ≤ 0.13σ, 판정 변화 ≤ 0.5% (≤ 5/1024) |
| 참고 크기 |
BF16 eager vs torch.compile 0.130σ, BF16 시드만 바꿀 때 0.392σ |
| 평가 경로 |
실행 |
기준 |
대상 |
| fake quant |
FP8 코드를 dequantize한 가중치 + 입력 FP8 quantize-dequantize (AQ FAKE_QUANT 평가와 비트 동일) |
FP adapter의 BF16 출력 |
22개 전부 |
| 서빙 |
NVIDIA cosmos-framework 정책 서버, TorchAO FP8 커널 (2 × RTX 4090, FSDP2 + CFG 병렬) |
같은 서버의 BF16 출력 |
5개 |
2. 서빙 결과
전달본 3개와, 전달 후 추가한 NVIDIA 형식 build 8개 중 fake quant 오차가 가장 낮은 SQ auto-α + GPTQ, 판정 변화가 가장 적은 SQ auto-α + AutoRound를 서빙했습니다. 다섯 build 모두 서버가 FP8 가중치 504/504개를 TorchAO로 로드했습니다.
| build |
서빙 σ [95% CI] |
서빙 RTN 대비 [95% CI] |
서빙 판정 변화 |
fake quant σ |
fake quant 판정 변화 |
gate (fake quant / 서빙) |
| RTN (전달) |
0.127 [0.095, 0.156] |
- |
12 |
0.097 |
9 |
초과 / 초과 |
| RTN + 이동평균 observer 0.01 (전달) |
0.099 [0.083, 0.117] |
-0.028 [-0.051, -0.004] |
7 |
0.105 |
4 |
통과 / 초과 |
| SQ auto-α + GPTQ |
0.104 [0.082, 0.127] |
-0.023 [-0.049, +0.003] |
8 |
0.098 |
8 |
초과 / 초과 |
| SQ auto-α (전달) |
0.106 [0.079, 0.137] |
-0.021 [-0.047, +0.005] |
4 |
0.100 |
4 |
통과 / 통과 |
| SQ auto-α + AutoRound |
0.116 [0.084, 0.149] |
-0.011 [-0.034, +0.010] |
5 |
0.109 |
2 |
통과 / 통과 |
차이가 어디서 나오는지 관측별로 나눠 봤습니다.
| build |
서빙 RTN 대비 |
RTN 대비 차이를 가장 키운 관측 2개 |
그 2개를 뺀 RTN 대비 |
| RTN + 이동평균 observer 0.01 (전달) |
-0.028 |
28, 2 |
-0.012 |
| SQ auto-α (전달) |
-0.021 |
5, 28 |
-0.006 |
| SQ auto-α + GPTQ |
-0.023 |
5, 28 |
-0.005 |
| SQ auto-α + AutoRound |
-0.011 |
28, 8 |
-0.000 |
- 네 build 모두 관측 28번이 차이를 키운 관측 2개에 들어 있습니다.
- 서빙 RTN은 오차 상위 관측 3개(5, 28, 2)가 제곱오차의 46%를 차지합니다. 같은 체크포인트를 fake quant로 돌리면 상위 관측이 7, 12, 15번으로 바뀝니다.
- 이동평균 observer는 95% CI가 0을 배제합니다. 다만 4개를 동시에 비교한 것을 Bonferroni로 보정하면 p ≈ 0.09입니다.
3. fake quant 결과
3.1 NVIDIA 형식 16개
| 순위 |
build |
σ [95% CI] |
RTN 대비 [95% CI] |
판정 변화 /1024 |
| 1 |
RTN (전달) |
0.097 [0.079, 0.116] |
- |
9 |
| 2 |
SQ auto-α + GPTQ |
0.098 [0.078, 0.120] |
+0.001 [-0.022, +0.025] |
8 |
| 3 |
SQ auto-α (전달) |
0.100 [0.078, 0.122] |
+0.002 [-0.023, +0.028] |
4 |
| 4 |
SQ auto-α, norm-FC 쌍만 |
0.103 [0.078, 0.131] |
+0.006 [-0.021, +0.035] |
8 |
| 5 |
SQ α=0.5 |
0.103 [0.075, 0.136] |
+0.006 [-0.024, +0.040] |
12 |
| 6 |
RTN + 이동평균 observer 0.05 |
0.104 [0.082, 0.127] |
+0.007 [-0.014, +0.031] |
9 |
| 7 |
RTN + 이동평균 observer 0.01 (전달) |
0.105 [0.080, 0.131] |
+0.008 [-0.016, +0.033] |
4 |
| 8 |
GPTQ |
0.106 [0.078, 0.133] |
+0.009 [-0.020, +0.037] |
7 |
| 9 |
SQ α=0.3 |
0.108 [0.084, 0.132] |
+0.011 [-0.009, +0.033] |
7 |
| 10 |
SQ auto-α + AutoRound |
0.109 [0.082, 0.136] |
+0.012 [-0.010, +0.036] |
2 |
| 11 |
AWQ + GPTQ |
0.111 [0.085, 0.136] |
+0.014 [-0.006, +0.037] |
12 |
| 12 |
SQ auto-α + 이동평균 observer 0.01 |
0.119 [0.087, 0.152] |
+0.022 [-0.012, +0.057] |
7 |
| 13 |
AutoRound |
0.121 [0.083, 0.160] |
+0.024 [-0.015, +0.065] |
8 |
| 14 |
RTN + RANGE_SHRINK |
0.132 [0.095, 0.166] |
+0.035 [+0.001, +0.068] |
21 |
| 15 |
AWQ |
0.135 [0.092, 0.175] |
+0.038 [-0.003, +0.077] |
9 |
| 16 |
RTN + HISTOGRAM observer |
0.212 [0.185, 0.240] |
+0.115 [+0.085, +0.143] |
17 |
- NVIDIA 형식 중 RTN보다 오차가 낮은 build는 없습니다.
- AutoRound는 AQ 구현상 INT, UINT, NVFP4 가중치만 튜닝합니다. 그래서 "AutoRound"와 "SQ auto-α + AutoRound" build의 FP8 코드는 각각 RTN, SQ auto-α와 바이트 단위로 같습니다(139억 개 중 0개 차이). 두 행이 RTN, SQ auto-α와 다른 것은 AutoRound 단계 뒤 activation calibration이 달라졌기 때문입니다(입력 scale 504개 중 250개, 302개). 즉 이 두 행은 AutoRound의 효과가 아닙니다.
- HISTOGRAM observer는 확실히 나쁩니다. RANGE_SHRINK는 보정 전 기준으로만 나쁩니다(21개 비교 보정 시 유의하지 않음).
- GPTQ는 AQ schema가 per-tensor weight를 막습니다. 실험 브랜치에서 이 검사를 풀었고, 원소 1개짜리 scale을 행 수만큼 펴는 수정도 넣었습니다. 이 수정이 없으면
Scale shape torch.Size([1, 1]) is incompatible with input shape torch.Size([4096, 1])로 실패합니다.
3.2 per-channel weight / per-token activation 6개 (탐색용)
RTN 대비 차이는 NVIDIA 형식 RTN(per-tensor)과 비교한 값입니다.
| 순위 |
build |
σ [95% CI] |
RTN 대비 [95% CI] |
판정 변화 /1024 |
| 1 |
SQ auto-α + GPTQ |
0.088 [0.066, 0.115] |
-0.009 [-0.037, +0.021] |
7 |
| 2 |
RTN |
0.090 [0.069, 0.112] |
-0.008 [-0.019, +0.003] |
8 |
| 3 |
AutoRound |
0.091 [0.073, 0.108] |
-0.006 [-0.026, +0.012] |
8 |
| 4 |
GPTQ |
0.095 [0.067, 0.128] |
-0.002 [-0.032, +0.033] |
4 |
| 5 |
AWQ |
0.096 [0.076, 0.114] |
-0.002 [-0.014, +0.011] |
5 |
| 6 |
SQ auto-α |
0.103 [0.078, 0.128] |
+0.006 [-0.014, +0.027] |
17 |
이 계열은 두 가지 이유로 지금 형식으로는 전달할 수 없습니다.
- NVIDIA 서버의 ModelOpt 체크포인트 로더(
apply_modelopt_fp8_checkpoint_inplace)는 두 scale을 (1, 1)로 읽어 PerTensor()로만 설치합니다.
- 같은 scheme은 프레임워크 런타임 FP8(
quantization.method=fp8, fp8_granularity=per_row)에만 있습니다. 이 경로는 BF16 가중치를 로드할 때 RTN으로 양자화하므로, AQ 알고리즘 결과를 담을 수 없습니다.
4. 결정과 남은 일
| 항목 |
내용 |
| 전달본 |
RTN, SQ auto-α, 이동평균 observer 유지. 두 경로 모두 gate를 통과한 전달본은 SQ auto-α뿐이고, RTN은 NVIDIA와 비트 동일한 기준으로 둠 |
| SQ auto-α + AutoRound |
두 경로 모두 gate 통과. 하지만 SQ auto-α보다 오차가 낮지 않아 추가하지 않음. 가중치는 SQ auto-α와 같음(AutoRound는 FP8을 튜닝하지 않음) |
| SQ auto-α + GPTQ |
두 경로 모두 판정 변화 8로 gate 초과 |
| 순위 확정에 필요한 것 |
calibration과 lab이 겹치지 않는 관측으로 다시 평가. /workspace/cosmos3-policy-quant/splits/splits.json(내부 서버)의 selection_validation(AUTOLab, IPRL, PennPAL, RPL)을 쓰면 됨. 지금 32개는 전부 calibration lab(TRI, CLVR, IRIS, ILIAD, RAD) 소속 |
5. 제약
- 평가가 in-sample입니다. calibration, 알고리즘 데이터, 선정이 모두 같은 32개 관측에서 나왔습니다.
- AQ는 DIFFUSION 모델에 SVDQuant만 허용합니다. 비교는 이 게이트를 푼 실험 브랜치
cosmos3-algo-sweep(미푸시)에서 했습니다. 커밋은 a3a1b56c4(게이트), ea30283f3(GPTQ per-tensor 검사), 8cdadc414(scale shape)입니다.
- vLLM-Omni 서빙은 해보지 않았습니다.
6. 재현 (내부 GPU 서버 경로)
| 항목 |
위치 (/workspace/cosmos3-policy-quant/) |
| config |
fq/configs/algo/*.yaml, fq/configs/algo3/*.yaml |
| 실행 스크립트 |
fq/algo_sweep*_run.sh, fq/algo_served4_*_run.sh |
| 순위표 |
fq/rank_all.sh → fq/eval/sweep_all.md, fq/rank_served.sh → fq/eval/served_all.md |
| 평가 하네스 |
harness/fakequant_loader.py, eval_fakequant.py, server_fidelity.py, stats_table.py, sweep_table.py |
Use
Served exactly as the RTN build: cosmos-framework's RoboLab action policy server with checkpoint_path set to this
repository and mixed_precision_first_steps=0, mixed_precision_last_steps=0, mixed_precision_w8a16_cache="none"
(generation under torch.no_grad, see the RTN build's card). fakequant_loader.py evaluates it as a fake-quantized
model on any GPU.
Caveats
- In-sample twice over: the calibration calls and the selection both use the 32 scored observations, so the selected
builds' margin over the others is optimistic.
- vLLM-Omni was not run; its step policy would serve W8A16 on the policy's 4 denoising steps.
Provenance
|
|
| Source |
nvidia/Cosmos3-Nano-Policy-DROID@805c0d6d46196ecdc789213a6de72303c383895b (BF16) |
| Quantizer |
np-quantizer-v2 experiment branch cosmos3-algo-sweep commit a3a1b56c4 (integration branch 411651222 + the DIFFUSION algorithm gate lifted), MODELOPT_FP8 in the diffusers layout; config in verification/config.yaml |
| ModelOpt state files |
nvidia-modelopt 0.44.0 |
License
A quantized derivative of nvidia/Cosmos3-Nano-Policy-DROID, distributed under the same
OpenMDW License 1.1. See NVIDIA's model card for intended use and limitations.