2953 shaares
3 results
tagged
inference
DFlash 2 decodes at close to 3× the speed of autoregressive decoding, about a third of the compute per token, with the same output.
Qwen3.5 Small models disable thinking by default. Use llama-server to enable it.