VYPR

Vllm

by Vllm

pypi: vllm

Source repositories

CVEs (88)

  • CVE-2025-46560MedApr 30, 2025
    risk 0.35cvss 6.5epss 0.01

    vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. Versions starting from 0.8.0 and prior to 0.8.5 are affected by a critical performance vulnerability in the input preprocessing logic of the multimodal tokenizer. The code dynamically replaces…

  • CVE-2025-29770MedMar 19, 2025
    risk 0.35cvss 6.5epss 0.00

    vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. The outlines library is one of the backends used by vLLM to support structured output (a.k.a. guided decoding). Outlines provides an optional cache for its compiled grammars on the local…

  • CVE-2026-73557MedAug 13, 2026
    risk 0.34cvss —epss 0.00

    vLLM is an inference and serving engine for large language models. From 0.20.2rc0 until 0.26.0, safe_load_prompt_embeds in vllm/renderers/embed_utils.py uses torch.sparse.check_sparse_tensor_invariants, whose process-global save, enable, and restore state can be raced by…

  • CVE-2026-9540MedMay 26, 2026
    risk 0.34cvss 5.3epss 0.00

    A vulnerability was identified in vllm-project vllm 0.19.0. This issue affects some unknown processing of the component OpenAI-compatible Serving Path. Such manipulation leads to denial of service. It is possible to launch the attack remotely. The exploit is publicly available…

  • CVE-2026-90554MedSep 12, 2026
    risk 0.33cvss 6.2epss 0.00

    vLLM versions >=0.10.2 and <0.28.0 do not apply any audio decode-size or duration limit when extracting audio from video input for NanoNemotronVL models. In nano_nemotron_vl.py, _extract_audio_from_videos calls load_audio_pyav(BytesIO(video_bytes)) without the max_duration_s or…

  • CVE-2026-34760MedApr 2, 2026
    risk 0.31cvss 5.9epss 0.00

    vLLM is an inference and serving engine for large language models (LLMs). From version 0.5.5 to before version 0.18.0, Librosa defaults to using numpy.mean for mono downmixing (to_mono), while the international standard ITU-R BS.775-4 specifies a weighted downmixing algorithm.…

  • CVE-2026-7141MedApr 27, 2026
    risk 0.29cvss 5.6epss 0.00

    A vulnerability was found in vLLM up to 0.19.0. The affected element is the function has_mamba_layers of the file vllm/v1/kv_cache_interface.py of the component KV Block Handler. Performing a manipulation results in uninitialized resource. It is possible to initiate the attack…

  • CVE-2026-54236MedJun 22, 2026
    risk 0.28cvss 5.3epss 0.01

    vLLM is an inference and serving engine for large language models (LLMs). Prior to 0.23.1rc0, the fix for CVE-2026-22778, which introduced a sanitize_message helper that strips object-repr memory addresses from error messages before they reach the client, is incomplete: several…

  • CVE-2026-34753MedApr 6, 2026
    risk 0.28cvss 5.4epss 0.00

    vLLM is an inference and serving engine for large language models (LLMs). From 0.16.0 to before 0.19.0, a server-side request forgery (SSRF) vulnerability in download_bytes_from_url allows any actor who can control batch input JSON to make the vLLM batch runner issue arbitrary…

  • CVE-2026-94625MedSep 21, 2026
    risk 0.27cvss 5.3epss —

    vLLM through 0.29.0 contains a resource exhaustion vulnerability in MooncakeConnector where rejected prefill requests create ownerless transfer placeholders that are never reclaimed. Attackers can send rejected requests to exhaust sender task pools, causing valid requests to be…

  • CVE-2026-92220MedSep 16, 2026
    risk 0.27cvss 5.3epss 0.01

    A vulnerability was found in vllm-project vLLM 0.26.0/0.27.0. Affected is the function MoRIIOConnectorScheduler.request_finished/MoRIIOConnectorWorker.get_finished/MoRIIOWrapper._handle_release_message of the file vllm/distributed/kv_transfer/kv_connector/v1/moriio/moriio_connect…

  • CVE-2026-78684MedAug 25, 2026
    risk 0.27cvss 5.3epss 0.00

    vLLM before 0.27.0 fails to properly classify DeepStream as a GPU backend and omits pixel-limit enforcement in its decode path. Unauthenticated attackers can activate DeepStream at request time to initialize the process-wide GPU decode pool and submit video that bypasses…

  • CVE-2026-73558MedAug 13, 2026
    risk 0.27cvss 5.3epss 0.00

    vLLM is an inference and serving engine for large language models. Prior to 0.27.0, an integer overflow in blockIdx.x * 2 * d in activation_kernels.cu can cause act_and_mul_kernel to consume another batched user's input, allowing a request processed in the same inference batch…

  • CVE-2026-73556MedAug 13, 2026
    risk 0.27cvss 5.3epss 0.00

    vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the structured_outputs.regex parameter in vllm/v1/structured_output/backend_lm_format_enforcer.py is passed to lmformatenforcer.RegexParser without compile_regex_with_timeout or validation in…

  • CVE-2026-73555MedAug 13, 2026
    risk 0.27cvss 5.3epss 0.00

    vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the validation_exception_handler in vllm/entrypoints/openai/server_utils.py converts FastAPI RequestValidationError objects with str(exc), and sanitize_message in vllm/entrypoints/utils.py does…

  • CVE-2026-12491MedJun 17, 2026
    risk 0.24cvss 4.8epss 0.00

    A flaw was found in vLLM, an open-source library for large language model inference. This vulnerability arises from improper handling of image metadata, specifically EXIF orientation and PNG transparency (tRNS) data, during image processing. When images are converted to RGB,…

  • CVE-2026-92365MedSep 16, 2026
    risk 0.21cvss 4.3epss 0.00

    A vulnerability was found in vllm-project vllm up to 0.29.0. Affected by this issue is some unknown functionality of the file vllm/v1/sample/thinking_budget_state.py. The manipulation results in inefficient algorithmic complexity. It is possible to launch the attack remotely.…

  • CVE-2026-90878MedSep 15, 2026
    risk 0.21cvss 4.3epss 0.00

    A vulnerability was determined in vllm-project vLLM up to 0.27.1. This affects an unknown part of the file /v1/chat/completions of the component Jinja Template Rendering. This manipulation of the argument chat_template causes resource consumption. The attack can be initiated…

  • CVE-2026-71486MedAug 17, 2026
    risk 0.21cvss 4.3epss 0.00

    vLLM is an inference and serving engine for large language models. Prior to 0.26.0, the /v1/completions/derender and /v1/chat/completions/derender endpoints accept caller-supplied GenerateResponse objects whose generate_responses, choices, token_ids, prompt_logprobs,…

  • CVE-2025-71379MedJun 20, 2026
    risk 0.21cvss 4.3epss 0.00

    vLLM versions >= 0.6.3 and < 0.9.0 contain multiple regular expression denial of service (ReDoS) vulnerabilities. Several regex patterns — in vllm/lora/utils.py, the phi4mini tool parser, and the OpenAI-compatible serving chat endpoint — are susceptible to catastrophic…

Page 4 of 5