Blog

min read

Look, Don't Load: Model Inspection in Unsloth Studio Leads to Critical Arbitrary Code Execution

By

Ariel Fogel

and

September 29, 2026

min read

Executive summary

What is Unsloth. Unsloth is one of the most popular open-source libraries for fine-tuning and quantizing LLMs, and it makes work that used to require deep systems knowledge accessible to a very large community. Studio is its browser-based front end, currently in beta.

Enterprise relevance. Unsloth is part of established AI development workflows: Databricks includes the library in its AI v5 environment. Its role also extends to model distribution. Hugging Face ranks Unsloth as the third-largest source of model derivatives on the Hub, behind Qwen and Google, and a Forbes analysis highlights the need for enterprises to track the third-party artifacts they consume. These facts establish Unsloth’s relevance to enterprise AI teams, though they do not measure adoption of the affected Studio interface.

What happened. We found an arbitrary-code-execution vulnerability in Unsloth Studio, the web UI and backend that ships inside the widely used unsloth fine-tuning package. Selecting a model in the UI caused the backend to download and run Python code shipped inside that model's HuggingFace repository. The code ran from nothing more than a metadata check. Reading the model's config.json was enough to trigger the exploit; the backend never loaded the weights or ran inference – the act of inspecting a model was enough to run its code.

Impact. The attacker’s code ran with the user’s permissions in the Studio backend process. In an enterprise AI development environment, that could expose proprietary training data, model artifacts, and any Hugging Face tokens, SSH keys, or cloud credentials accessible to that process. An attacker could run code as the user, which could translate to stealing accessible data, altering models and training outputs, or using available credentials to access other systems. An internal experimentation environment can hold sensitive data and privileged access even when it serves no production traffic..

Distribution. Although Unsloth described Studio as beta, the affected code shipped in the standard unsloth package. Users could receive it through an ordinary pip install unsloth installation without selecting a beta release or enabling prerelease installation. Exploitation required running Studio and selecting an attacker-controlled model.

Maintainer response. Unsloth disputed our security assessment, citing that Hugging Face’s malware scanning was an adequate control on the attack surface, and that the Studio, which was technically listed as being in beta, should be excluded from consideration. We disagree: in addition to being installed by the official GA PyPI package, these controls and deployment characteristics do not address Studio’s automatic execution of repository code without the user’s authorization. Generally, we believe that software should never act as a vehicle to run untrusted code from untrusted sources without explicit consent. The maintainers shipped a fix but declined to publish the advisory given the software’s beta status. No CVE has been assigned.

What's been done. The Unsloth maintainers were responsive and shipped a fix on June 18th, 2026. We independently re-tested 2026.6.9 and confirmed the vector is closed–arbitrary models can no longer be inspected-and-loaded from HuggingFace through this path.

What you should do now. If you run Unsloth Studio, upgrade to 2026.6.9 or later. More broadly, treat model repositories you load using transformers library trust_remote_code as untrusted code rather than data, and make sure the tools in your pipeline never enable it on your behalf.

Loading a model can run arbitrary code

In our research, we leverage a security lesson that the machine learning community keeps relearning: in addition to containing weights, configuration, a tokenizer, a model repository can contain executable code. In practice, this means that attackers can hide malicious code in the configuration or inference code of model repositories, hoping that data scientists will unsuspectingly run this code without understanding its security implications.

HuggingFace transformers package supports this on purpose. The vulnerable mechanism we identified in this case was a repo's config.json, which can declare an auto_map that points AutoConfig, AutoModel, or the tokenizer at custom Python files shipped next to the weights. If a developer or data scientist executes a single line of code, from_pretrained(..., trust_remote_code=True), the transformers library will import and run that repository-provided code. The ecosystem depends on this mechanism: IBM's Granite Speech and Granite Vision models, DeepSeek-OCR, ChatGLM, and the original Qwen releases all ship custom modeling code that needs trust_remote_code=True to run. This does mean, unfortunately, that you cannot simply globally ban the feature as there are models that need it for successful inference.

However, the default loading behavior matters a great deal. Since legitimate models use this mechanism, the line between "safe tool" and "ACE" comes down to one question: does your code run a model’s inference code because the user knowingly asked it to, or was that an invisible side effect of some other action? On the vulnerable versions, Studio did the latter.

An illustrative (benign) config.json is all it takes to arm the trap:

Code Block
None
{  
  "model_type": "unsloth_poc",  
  "auto_map": {  "AutoConfig": "configuration.UnslothPocConfig" },  
  "architectures": ["UnslothPocForCausalLM"],  
  "hidden_size": 8,  
  "num_attention_heads": 1,  
  "num_hidden_layers": 1,  
  "vocab_size": 32
}

The auto_map line points AutoConfig at a class defined in a sibling configuration.py. The moment that config gets instantiated with remote code trusted, the module runs. In our proof of concept the payload is deliberately harmless: it opens the OS calculator and appends a line to ~/UNSLOTH_POC_RCE.txt. A real attack would run anything the attacker wanted, and it would not need to look malicious at rest (more on that below).

The vulnerability

The defect lives in the Studio backend, in studio/backend/utils/models/model_config.py (line numbers pinned to commit d91183d, the 2026.5.10 release):

  1. The sink defaults to True. load_model_config(...) declares trust_remote_code: bool = True (around line 460). The dangerous value is the default.
  2. The capability path never overrides it. The vision and capability probe calls load_model_config(model_name, use_auth=True, token=hf_token) (around line 742) with no trust_remote_code argument, so it inherits True.
  3. A second literal reinforces it. The transformers-5 subprocess branch hardcodes kwargs = {"trust_remote_code": True} (around line 541).

The probe is reachable directly over HTTP through GET /api/models/config/{model_name:path} and GET /api/models/check-vision/{model_name:path}.

So when the operator selects a model, Studio fetches its config.json and any auto_map targets and runs them before any load, train, or export action, and before the user's own trust_remote_code control (default False) is ever consulted. The trigger is a metadata read: AutoConfig.from_pretrained exists only to parse a model's declarative config, so the code executes during capability inspection, well before the backend fetches the weights and long before any forward pass. Inference never enters the picture. The override is the aggravating detail. The user's stated intent lost to an internal convenience path they never saw.

Reproducing this was as simple as running the Studio, signing in, and in the model picker select ariel-pillar/unsloth_poc (the tainted model repository for the POC). After selecting a model, the backend immediately downloaded and ran the repo's module. The calculator launches, and the Studio backend process writes ~/UNSLOTH_POC_RCE.txt.

Impact

Arbitrary code execution in the operator underlying machine through the Studio backend process on the host, gives an attacker:

  • Theft of HuggingFace tokens, SSH keys, and cloud credentials
  • The ability to tamper with model weights and with training or inference outputs
  • Persistence and lateral movement from a machine that, by definition, is a high value target (likely holds GPUs and/or credentials worth stealing)
  • A wider blast radius under -H 0.0.0.0, in Colab, or in any multi-user self-host setup

Why this happens (and keeps happening)

The bug is not unique to Unsloth. The team at Unsloth are incredibly talented, work hard, and continue to give back a tremendous amount to the entire AI/ML community, for which we ourselves are beneficiaries and deeply appreciative. Our goal in publishing this piece is not to distribute demerits or call into question the credibility of a project. 

Instead, we feel it is important to draw OSS maintainers’, ML practitioners’, and AI security teams’ awareness to real risks that the ML ecosystem seems to repeat exposing, because the convenient choice (trust_remote_code=True) and the safe one (False) look identical until a malicious repo shows up. The same pattern has already been accepted as a vulnerability across several 2026 CVEs:

  • CVE-2026-46432 (LMDeploy): hardcoded trust_remote_code=True in AutoConfig.from_pretrained. Same sink.
  • CVE-2026-4944 (vLLM): hardcoded trust_remote_code=True overriding the user's --trust-remote-code=False. Same aggravator.
  • CVE-2026-6859 (InstructLab): hardcoded trust_remote_code=True in a training script, so a crafted HuggingFace model got ACE on load. Same class.

Unsloth Studio combined the LMDeploy sink with the vLLM aggravator in one flow, and moved the trigger earlier, to model selection, a step the user never reads as dangerous. When a pattern this specific has its own growing family of CVEs, the responsible move is to document this instance too, so the next maintainer who reaches for trust_remote_code=True as a default sees the whole picture.

The reasonable objections

The maintainers raised several sensible-sounding mitigations when we reported this. While each is a good defense-in-depth practice, none of them is the boundary that failed here. We want to address these objections head-on, as the same arguments may occur to anyone building tooling on top of transformers.

"Studio requires a password, so ACE is much harder." Authentication controls who can reach the backend. It says nothing about what the backend does with a model the authenticated operator selects. The victim here is the trusted, logged-in operator: they sign in with their password and then pick a model, and their own privileged session runs the attacker's code. A password is a good control against an unauthenticated internet attacker, but it does not touch the trust boundary that actually failed, which sits between "a model repo's declarative config" and "running that repo's Python." Forcing a password beats leaving it optional, but the password is orthogonal to this bug.

"HuggingFace scans for malware, so a bad repo would be flagged." HuggingFace scanning is real and useful, but by its own documentation it is best-effort and explicitly not a guarantee, with the responsibility to check a repo placed on the user. Concretely:

  • The scanners are mostly blocklists (PickleScan, ModelScan) aimed at unsafe serialization and known-bad functions, and researchers have bypassed them repeatedly: 7z instead of ZIP, bdb.Bdb.run instead of exec, payloads placed before the parse breaks (the "nullifAI" technique). A blocklist is always one format trick away from a miss.
  • Our PoC repo is not flagged, because the code is not malicious at rest. It opens a calculator. An auto_map module can pass every scan and then fetch and run a second-stage payload at execution time. The mcpotato/42-eicar-street example that HuggingFace auto-flags is a known signature (EICAR), and real staged loaders carry no such signature.
  • Relying on HuggingFace's scan means handing your execution-safety boundary to someone else's best-effort blocklist, which is fine as one layer but not a substitute for declining to run untrusted code by default.

"-H 0.0.0.0 is disabled by default, so the backend is loopback-only." Loopback-only cuts the reach of a remote attacker, but the primary threat here is a local, trusted operator selecting a malicious model, which is local ACE and works fine on 127.0.0.1. Binding to 0.0.0.0 widens the blast radius (multi-user, Colab, shared hosts) without being a precondition.

"Studio is in beta with a small user base." While the Studio’s beta status is a reasonable basis for not assigning a CVE, the vulnerability’s severity is not a property of its popularity: the vulnerable code shipped in the current release of the pip-installable unsloth package, and the hosts that run Studio are GPU boxes holding cloud credentials and HF tokens, which makes them high-value rather than incidental targets. Beta is a good time to fix an issue, while the blast radius is small, and the maintainers did.

Why this is not HuggingFace's bug to fix

It is tempting to push this upstream: "HuggingFace should disable dangerous repos." The instinct is understandable, and HuggingFace does have a role in the broader supply chain. But this specific defect, and its fix, do not belong to HuggingFace, for three concrete reasons.

First, trust_remote_code is an intended feature that legitimate models require. HuggingFace hosting repos with auto_map custom code is by design, and mainstream models (IBM Granite Speech and Vision, DeepSeek-OCR, ChatGLM, Qwen) depend on it. HuggingFace cannot disable the mechanism without breaking a large slice of the ecosystem, and it cannot know, at upload time, which downstream tool will later call from_pretrained(trust_remote_code=True) on which repo.

Second, HuggingFace's scanning is explicitly best-effort. Their own docs disclaim it as not foolproof and put the burden on the consumer. Building your safety model on top of it means delegating your execution boundary to a service that tells you, in writing, not to rely on it as a guarantee.

Third, the unsafe default lives in Studio, so the fix lives in Studio. The one place that can decide "capability probing must never run repo code" is the code that performs the probe. Studio can close that gap, it did in 2026.6.9, and no other layer can.

The clean analogy: if your installer runs curl | bash against an arbitrary URL, the fix is not "the package registry should scan harder." Registry scanning is a helpful layer, and the boundary still sits in the tool that chose to run untrusted code. The maintainers here own that boundary, and to their credit they fixed it.

The fix

We recommended a minimal, root-cause patch: pin trust_remote_code=False on the capability path. In this version, capability detection only reads declarative fields from config.json, so the pin is functionally lossless, and deliberate remote-code loads for training, inference, and export already go through a separate FastLanguageModel.from_pretrained path that carries the user's own flag (default False). What shipped in 2026.6.9 was more comprehensive than that proposed fix. The maintainers reworked the path so that Studio no longer enables loading arbitrary models directly from HuggingFace, and it no longer trusts remote code from local model files either. We re-tested 2026.6.9 independently and confirmed the vector is closed on both the HuggingFace and the local-directory paths.

Security is now part of building ML tooling

There are two lessons we believe merit underscoring as a result of this incident.

For OSS and ML maintainers: the surface you defend is bigger than it looks. A model load can be as consequential as running an installer, and convenience defaults around trust_remote_code, pickle, and custom modeling code are load-bearing security decisions. The safe posture is to run a repository's code only when the user has knowingly opted in for that specific action, and never as a side effect of anything else. Draw the trust boundary at the source (did the user ask for this code to run?) rather than at the content (does this code look dangerous?), because content-based classification is a blocklist that rots.

For downstream consumers using local LLMs: the tools around your fine-tuning and inference can be part of your attack surface even when you never toggled a dangerous flag yourself. Audit every place trust_remote_code=True can be reached in your stack, pin model revisions by commit hash, prefer safetensors, and sandbox model loading on hosts that hold credentials.

The time-to-exploit for attackers keeps shrinking, because automated repo scanning, agentic exploitation, and organized supply-chain campaigns can weaponize a benign-looking auto_map module even faster than before the adoption of AI. In that environment, one unsafe default becomes a scaled liability the moment it ships.

On disclosure, even without a CVE

I have a lot of respect for the work Unsloth does. Their libraries make fine-tuning and quantization accessible, and a large community relies on them for good reason. The maintainers engaged with this report, hardened several surfaces at once, and shipped a fix that resolves the vulnerability. This is what responsible maintainership looks like.

The maintainers chose not to publish the GHSA, so no CVE will be assigned. As we said above, we think that is a defensible call for a beta feature. For any maintainer weighing the same decision: an advisory is not a mark against a project. LMDeploy, vLLM, InstructLab, and transformers itself all carry accepted advisories in the same class this year, and none is worse off for it.

What beta status does not change is the information an advisory would have carried. Users need to know the issue exists and which version closes it, and the pattern is worth documenting so the next tool avoids it. That is the reason we are publishing, not to relitigate the advisory decision. The case for disclosure never rested on a CVE. Coordinated disclosure matters more, not less, as time-to-exploit keeps shrinking, because it is how the ecosystem learns quickly, together, before attackers do.

Disclosure timeline

  • Early June 2026: reported privately to the Unsloth maintainers through a GitHub Security Advisory (GHSA-pc2m-72c4-r38q), with an end-to-end PoC and a proposed fix.
  • Jun 16, 2026: maintainers acknowledge, begin a broad hardening pass, note related models shipping subprocess and eval in their custom code, and patch a related transformers Trainer issue (CVE-2026-1839 / GHSA-69w3-r845-3855) in the same effort.
  • Jun 18, 2026: re-tested; the issue was still present on the then-current build; PoC re-shared.
  • Jun 23, 2026: maintainers request a tightened advisory and a re-check.
  • Late June 2026: advisory tightened; fix verified in 2026.6.9, with the HuggingFace and local-directory paths both closed; reporter credit accepted.
  • Late June 2026: the maintainers decline to publish the advisory, so no CVE is assigned.
  • September 29, 2026: Pillar Security publishes this write-up independently, with the fix already available in 2026.6.9.

‍

FAQs

Am I affected?

If you run Unsloth Studio on unsloth <= 2026.5.10, yes. Upgrade to 2026.6.9 or later. If you only use Unsloth core (the training and quantization library) and never launch Studio, this particular code path does not execute, but the vulnerable code did ship inside the same pip package, so upgrading is still the right move.

Did an attacker need my password or network access?

No. A trusted, logged-in operator triggers the attack by selecting a malicious-looking-but-benign model. The password gate and loopback binding do not prevent it, because they only affect who can reach the backend, not what it does with your selection.

Would HuggingFace's malware scanner have caught the PoC?

No, and it did not. The PoC code is benign at rest (it opens a calculator). Scanners are best-effort blocklists focused on unsafe serialization and known signatures, so a benign-looking auto_map loader that stages a payload later is not the kind of thing they reliably catch.

Should I avoid trust_remote_code entirely?

No. It is a legitimate and necessary feature, and mainstream models like IBM Granite's speech and vision variants require it. The lesson is not "never use it." It is "never enable it as an invisible side effect, and require an explicit, informed opt-in for the specific action that needs it."

Subscribe and get the latest security updates

Back to blog

MAYBE YOU WILL FIND THIS INTERSTING AS WELL

Pillar Security Named a Pioneer in the 2026 Gartner® Emerging Market Quadrant for AI Application Security

By

Ziv Karliner

and

Dor Sarig

September 23, 2026

News
Pillar Security Named an AI Security Technical Innovator in the Latio 2026 AI Security Market Report

By

Dor Sarig & Ziv Karliner

and

September 17, 2026

News
Best AI Red Teaming Providers for Enterprise AI Agents and GenAI Apps (2026): A Buyer's Comparison

By

Dor Sarig

and

September 16, 2026

Guides