MTP broken?
I'm getting almost 0% Avg Draft acceptance rate using this model whereas Pilcothink, GoldHub, and Minachist auto round quants all have normal acceptance rate with the same settings. I know the model card said --speculative-config works out of the box but do I need to do anything special to get it working?
(using vLLM v0.25.1)
--speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
Thanks for the report. I couldn't reproduce this: stock vLLM 0.25.1 (fresh pip install), this quant, TP2 on RTX 3090s, --speculative-config '{"method":"mtp","num_speculative_tokens":3}', no other quant flags β avg draft acceptance 54β58%, mean acceptance length ~2.7, both greedy and temp 1.0.
One relevant difference vs the other AutoRound quants you tried: this one also quantizes the MTP head to int4 (they keep it bf16), so a kernel-path difference on your side could matter. Could you share: GPU model, TP size, full launch command, and the exact SpecDecoding metrics log line you're seeing? That would help pin it down.
Thank you for confirming. I obviously had something setup incorrectly. I downloaded the quant to try again and for whatever reason MTP is working this time (must have been user error).
With simple snake game creation test I'm seeing...
Avg Generation Speed: ~96.7 tokens/s
Mean Acceptance Length: 3.24 tokens
Overall Draft Acceptance: ~74.3%
Per-Position Acceptance Rates:
Position 1: 84.7%
Position 2: 73.3%
Position 3: 65.6%