Unsloth Studio Flaw Allows Malicious AI Models to Execute Code
A vulnerability in Unsloth Studio enabled malicious AI models to execute arbitrary Python code simply by being inspected, posing a risk to sensitive enterprise AI development environments.

A critical vulnerability discovered in Unsloth Studio, the web UI front end for the popular Unsloth library used in fine-tuning and quantizing large language models (LLMs), has been patched. The flaw allowed malicious AI models to execute arbitrary Python code on a user's system simply by being inspected, bypassing the need for inference or loading model weights.
Pillar Security's Ariel Fogel detailed the vulnerability, explaining that the issue stemmed from Unsloth Studio's use of the trust_remote_code=True setting when checking a model's configuration. This setting, intended to allow the underlying 'Transformers' library to download and execute custom Python code referenced in a model's config.json, was triggered during a routine metadata check. Consequently, the backend would execute the model's code without ever loading the model weights or performing inference, making the act of inspecting a model sufficient to trigger the exploit.
The potential impact of such an exploit is significant, particularly within enterprise AI development environments. An attacker could leverage this vulnerability to steal proprietary training data, compromise model artifacts, and exfiltrate credentials such as cloud logins or SSH keys tied to the compromised process. Fogel emphasized that even internal experimentation environments, which may not handle production traffic, can contain highly sensitive data and privileged access.
While Pillar Security has not observed any real-world exploitation or malicious model repositories specifically targeting this mechanism to date, Fogel noted that malicious models uploaded to platforms like Hugging Face have been used in other campaigns. The vulnerability could allow an attacker to execute code with the user's permissions, leading to data theft, model tampering, or unauthorized access to other systems.
Pillar Security reported the vulnerability to Unsloth in early June, and the issue was addressed in update 2026.6.9 later that month. Pillar confirmed the fix was effective. However, the security vendor noted that Unsloth disputed aspects of their security assessment, suggesting that Hugging Face's malware scanning was sufficient and that the Studio, being in beta, should be excluded from consideration. Pillar disagreed, arguing that the automatic execution of repository code by the Studio remained a significant risk.
No CVE identifier was assigned to the vulnerability, as Unsloth declined to publish a security advisory. Pillar Security recommends that users upgrade to Unsloth Studio version 2026.6.9 or later. More broadly, they advise treating model repositories loaded via the Transformers library with trust_remote_code enabled as untrusted code rather than mere data, and ensuring that development pipelines do not enable this setting automatically.
This incident is not isolated; similar vulnerabilities involving the trust_remote_code setting have been observed in other machine learning tools, including LMDeploy (CVE-2026-46432), vLLM (CVE-2026-4944), and InstructLab (CVE-2026-6859) earlier this year. The recurrence of these issues points to a systemic gap in how machine learning tools handle executable content within model artifacts.
Fogel highlighted that the Unsloth case is particularly concerning because the security boundary was crossed during an action that users would reasonably perceive as a simple inspection. This underscores the need for greater transparency and user control over how AI tools handle potentially executable content embedded within models.