Seminars
View all Seminars | Download ICal for this eventClosing the Heterogeneity Gap: Lossless, Training-Free Self-Speculative Inference of Large Language Models
Series: M.Tech (Research) Colloquium
Speaker: Viren Luke Radhakrishnan, M.Tech (Rssearch) ERP student, Dept. of CSA
Date/Time: Oct 08 14:15:00
Location: CSA Seminar Hall (Room No. 254, First Floor)
Faculty Advisor: Prof. Chiranjib Bhattacharyya & Dr. Prakash S Raghavend
Abstract:
Large language models (LLMs) are increasingly used for coding and reasoning, and there is growing interest in running open-weight models locally for cost, privacy, and offline use. Beyond roughly 24B parameters, however, these models no longer fit in the VRAM of a typical workstation GPU or single-GPU cloud instance. Layers that do not fit are offloaded to host DRAM and either streamed to the GPU on demand or executed on the CPU. Both limit throughput: streaming is bound by PCIe bandwidth, and CPU execution is too slow for interactive use. Speculative decoding is the dominant lossless acceleration framework in this regime: a small draft model proposes several tokens, which the full model verifies in a single forward pass, and rejection sampling guarantees an output distribution identical to the full models. Existing heterogeneous (CPUâ??GPU) methods, however, leave much of the hardware idle, a problem we call the heterogeneity gap. Only one device performs useful work at any instant, and the draft is a separate model, often trained per target at considerable cost, that occupies VRAM while sitting idle during verification. In this work, accepted as a poster to NeurIPS 2026, we present Speculative Tensor Parallelism (STeP), a lossless, training-free, sel f-speculative method that closes this gap. Channel-saliency methods from structured pruning order each FFN layers channels so that the top-k channels closely approximate the full layers output, degrading gracefully as k decreases. STeP exploits this by reordering rather than pruning the channels, placing the top-k on the GPU as the draft and the rest on the CPU, with k chosen to fill the available VRAM. During verification, the GPU and CPU compute their shards of each layer concurrently, and their sum recovers the full layers output exactly. The GPU-resident channels thus play a dual role: an independent subnetwork draft and a tensor-parallel shard of the verifier. We evaluate STeP on dense models from 24B to 123B parameters, GPU memory budgets from 32 to 192 GB, and five benchmarks spanning coding, reasoning, summarization, and chat. STeP outperforms two leading methods, SpecExec and SubSpec, by up to 1.6Ã? under greedy decoding and 2.2Ã? at temperature 0.6, and is faster than either of its components, tensor parallelism and speculative decoding, used alone.
Meeting Link :
Link
