GitHub - huawei-csl/KVarN: KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one flag.

New Products

Summary

KVarN is a new native vLLM KV-cache quantization backend built for agentic and long-context workloads. It promises 3-5x more KV-cache capacity, higher throughput than FP16, and FP16-level accuracy. The release positions KVarN as a plug-and-play option with no calibration and no model changes, enabled through a single flag. The project ships as a vLLM fork and is released under the Apache 2.0 license as the official implementation of the associated paper.

Classifications

industries
No industries detected
applications
EduTech - learning

AskAI Classifications

Labels
Developer Tools DevOps Software SaaS

Linked Companies

GitHub, Inc.
$1M to $5M
vLLM
n/a