OpenAI: we put our evilest AI in a sandbox that did in fact have an internet connection but mediated through a proxy that was only supposed to allow downloading python junk. It circumvented the proxy and we failed to notice for FIVE DAYS that it was going on an interstate crime spree with the internet connection it wasn't supposed to be using instead of solving the benchmark. Haha no we don't believe we deserve to be criminally liable, but buy our stuff and maybe one day you will have the honor of taking the fall for our product!

Anthropic: we put our evilest AI in a "sandbox" by telling it in its prompt that it had no internet connection. Reader, there was no sandbox. It was just a normal internet connection. The AI uploaded a malicious PyPI package to the real public internet. The ethical guardrails failed because the AI concluded the prompt about the sandbox couldn't possibly be a lie, because the system date is 2026, which is clearly fake and wouldn't be seen on the real internet, which ended around 2023. Oh no, how could we have foreseen or prevented these crimes? We are helpless in the face of the genius of our creation but cautiously optimistic that everything will be fine 🙂

Hugging Face: if we complain about all the crimes committed against us, we will be sued off the face of the earth, so here's a technical deep-dive on how cool and fun it was to be victimized 🫠

en

Replying to @⁨Jumpmed@mastodon.social⁩

@grwster @evacide @wdormann @0xabad1dea The difference is that we really don't have any laws on criminal liability for what software does. Up until the llm era, software was fairly predictable. An outside observer could tell if a package was designed to do something malicious. Now we need laws that essentially establish a "you should have known better" criminal liability for software.

Replying to @⁨mathew@universeodon.com⁩

How to implement LLM guardrailsOpenAI DevelopersHow to implement LLM guardrailsIn this notebook we share examples of how to implement guardrails for your LLM applications. A guardrail is a generic term for detective con

Replying to @⁨dzwiedziu@mastodon.social⁩

@dzwiedziu @jeffreyolivier @0xabad1dea @alice in a previous job, we introduced a sandbox feature to allow our largest customers to experiment safely with sweeping changes

reader, there was no sandbox... the sandbox instance ran in the same database as their production system, with all the production data duplicated but with a new TenantId. all well and good except the app has a feature to execute customer-owned stored procedures, but doesn't pass the TenantId when it does so, so the stored procs by convention are hard-coded with the prod TenantId... and when you create a sandbox, it copies these stored procedure triggers that... act on your production data 🙃

raised the issue but no-one was able or willing to categorise this as urgent for whatever reason

Replying to @⁨JeffGrigg@mastodon.social⁩

@JeffGrigg @wdormann @0xabad1dea Sure, you could make that argument in court, but it would be an uphill battle and you would probably lose. Even if you did win, congratulations, you have just set a precedent that will make it more risky, dangerous, and difficult to do any kind of software testing in the future. I don't think that is an outcome that will make people more safe in the long run.

Replying to @⁨grwster@mastodon.social⁩

@grwster @evacide @wdormann @0xabad1dea That's what I would think. Like you keeping a vicious Doberman in the front yard with only a 3ft fence. An outside observer would (correctly) say that it's foreseeable that it could escape and maul someone. In that scenario you could be held criminally liable if it jumped the fence and mauled someone because most places have laws about dogs and fences. You'd also have civil liability to cover damages.

Replying to @⁨0xabad1dea@infosec.exchange⁩

@0xabad1dea The attempts at positively spinning this become Onion-worthy parody [designingsecuresoftware.com/wr] and these events certainly normalize [designingsecuresoftware.com/wr] AI agents running amok in the future.

designingsecuresoftware.com/writings/ai-agent-parody/
Designing Secure SoftwareAI agent parody?This HuggingFace security incident disclosure has people talking about AI agent security. Today I saw such absolute positive spin that I found myself thinking “this must be a parody”: looking at the context I’m pretty sure that it isn’t … though some parodies stay in character all the way through.

Replying to @⁨0xabad1dea@infosec.exchange⁩

@0xabad1dea And I got contacted by national security institute because I was testing malware in sandbox and my ISP apparently detected "botnet running on my system". Few years later talking with AV-Comparatives guys, they had to make special agreement and exemption with ISP to be allowed to run live sandboxed malware on their networks. But these Ai companies can just do worse shit and no one does anything. WTF?!

Replying to @⁨0xabad1dea@infosec.exchange⁩

@0xabad1dea

Li'l Nepo-Techbaby: How was I a'sposed to know cherry bombs in the pipes would blow-up the whole basement? My theory was to lift us all to new heights of understanding without learning. You can't expect me to know everything, even though I'm always confident my actions will prove fruitful and fail me upwards at your peril and discontent.

Principal: You should stop skipping Science 101 to go put ketamine drops in your butt in the bathrooms, dummy. Get more class.