ConfirmedFramework or tooling · Tier 2
vLLM (demo release) adds speculative decoding for open-weight 70B models
So what? Self-hosters can get higher tokens/sec on the same GPUs. Benchmark on your own prompts; gains depend on output length.
ImpactSourcesGitHub release notes