VYPR

Vllm

by Vllm

pypi: vllm

Source repositories

CVEs (88)

  • CVE-2025-46722MedMay 29, 2025
    risk 0.20cvss 4.2epss 0.00

    vLLM is an inference and serving engine for large language models (LLMs). In versions starting from 0.7.0 to before 0.9.0, in the file vllm/multimodal/hasher.py, the MultiModalHasher class has a security and data integrity issue in its image hashing method. Currently, it…

  • CVE-2025-61620medOct 7, 2025
    risk 0.19cvss —epss 0.00

    ### Summary A resource-exhaustion (denial-of-service) vulnerability exists in multiple endpoints of the OpenAI-Compatible Server due to the ability to specify Jinja templates via the `chat_template` and `chat_template_kwargs` parameters. If an attacker can supply these…

  • CVE-2026-93841LowSep 18, 2026
    risk 0.17cvss 3.7epss 0.00

    vLLM through 0.29.0 contains a memory corruption vulnerability in the Triton _bincount_kernel where prompt token IDs index the penalty prompt-presence bitset without bounds checking against vocabulary size. Attackers can submit multimodal audio requests with tokens equal to…

  • CVE-2026-93840LowSep 18, 2026
    risk 0.17cvss 3.7epss 0.00

    vLLM before 0.29.0 validates allowed_token_ids against tokenizer length instead of model output logits width in SamplingParams._validate_allowed_token_ids(). Attackers can supply token IDs above the output vocabulary that pass validation, causing LogitBiasState to corrupt GPU…

  • CVE-2026-90713LowSep 14, 2026
    risk 0.14cvss 3.3epss 0.00

    A security flaw has been discovered in vllm-project vLLM up to 0.29.0. The affected element is the function TiktokenTokenizer::new of the file rust/src/text/src/backend/hf/mod.rs of the component tiktoken vocab File Handler. The manipulation results in denial of service. The…

  • CVE-2026-93989LowSep 19, 2026
    risk 0.13cvss 3.1epss 0.00

    vLLM through 0.29.0 fails to properly validate bad_words token indices against the model's generation output width in SamplingParams.update_from_tokenizer(). Attackers can supply out-of-bounds token indices that corrupt logits memory of concurrent requests, causing different…

  • CVE-2025-46570LowMay 29, 2025
    risk 0.10cvss 2.6epss 0.00

    vLLM is an inference and serving engine for large language models (LLMs). Prior to version 0.9.0, when a new prompt is processed, if the PageAttention mechanism finds a matching prefix chunk, the prefill process speeds up, which is reflected in the TTFT (Time to First Token).…

  • CVE-2025-25183LowFeb 7, 2025
    risk 0.10cvss 2.6epss 0.00

    vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Maliciously constructed statements can lead to hash collisions, resulting in cache reuse, which can interfere with subsequent responses and cause unintended behavior. Prefix caching makes use…

Page 5 of 5