Log

Security

We assume your AI will be fooled eventually. What we have tested, what we have not, and why that distinction is printed.

VaultHelm assumes your AI will eventually be fooled. Not might be — will be. Everything here is built so that when it happens, it doesn’t matter much.

The thing we’re actually defending against

You can hide instructions inside almost anything an AI reads. A line buried in a document. Text in an email. A comment on a web page. The AI reads it as if you’d typed it, and does what it says.

This isn’t exotic or rare. The moment an AI can read your email, browse the web, and open your files all in the same conversation, it’s simply the weather. Any system that treats it as an unlikely edge case is one badly-worded PDF away from a bad afternoon.

So VaultHelm starts from the assumption that your AI can and will be talked into something at some point — and makes sure that being talked into it isn’t enough.

You can change an AI’s mind. You can’t change what it’s able to reach.

Why this isn’t just careful instructions

This is the part worth understanding, because it’s what separates VaultHelm from a well-written prompt.

Nobody tells the AI to behave. There’s no instruction saying “please ask before doing anything important,” because instructions are exactly what an attacker overwrites. Every limit lives on the server, outside the conversation, where the AI has no reach at all.

The practical test: someone sends you a document engineered to hijack your assistant. It works — the AI is fully convinced it should go and change something important. It then can’t, because the ability was never handed to it in the first place. The attacker won the argument and got nothing, because the argument was never what was holding the door shut.

And it can’t promote its own work

Your AI can always write — into a holding area. What it cannot do is move anything out of there and into the real record. That takes a separate step it has no way to reach, and that step only runs because you said so.

This is an old idea rather than a clever one: the person who writes the cheque doesn’t sign it. Keeping the two apart means a compromised assistant can fill the holding area with whatever it likes and still change nothing that counts.

Nothing is lost while you decide, either. Staged work sits there indefinitely and stays readable until you accept it, correct it, or bin it.

What it’s allowed to open, not just what it’s allowed to do

Most systems police reading with a list of file names that look like they hold secrets. We went looking for evidence that ours was good enough and found the opposite: a certificate authority’s private key, a server key, a client certificate bundle and a web server’s private key, all sitting in a backup someone had filed, all openable. The list held every obvious name and matched none of them — because files like that are named after the service they belong to, not after what’s inside them.

A list of bad names only protects the names somebody thought of. Files of the kind that hold keys and passwords are now refused before they are opened, whatever they happen to be called, and the handful of exceptions are ones we could prove we needed.

We didn’t ship that on a hunch, and we didn’t ship it on a test we wrote ourselves either. We took every file the system had genuinely opened in its life and put all of it back through the new rule before it went near a customer. Exactly one came back refused, and it was the one the rule was written for.

One bounded gap, stated plainly because you’d find it eventually: ordinary data files are still readable. Refusing those would mean an exceptions list longer than the rule it replaced, so they are swept separately instead, every time we ship a change.

Which of these we have actually proven

This page is written in the present tense — “can’t”, “is refused”, “won’t reach” — and you should not take the whole of it at the same weight. Some of these have been proven against the running system by breaking them deliberately and watching the check catch it. Others are live code that we have not yet tried hard enough to defeat, and at least one still owes a proper attack test.

The difference is printed rather than glossed. The tested page lists every claim below with a third column — proven, working, or owed — and it is the more honest of the two pages. Read it before you believe this one. If that ordering seems backwards for a security page, it is the whole point: a list of features with no status column is the thing we are trying not to sell you.

What’s actually in place

Concern What stops it
The AI changing something important It can’t write to anything that matters — its work lands in a holding area, and moving it into the real record takes a step the AI cannot reach and you have to take.
The AI reaching a tool it shouldn’t The list it can see and the list it can use are the same list, checked in one place. There’s no hidden menu.
One customer seeing another’s work Separate installations, separate files, separate keys — even though we host them. There is no shared pile.
Passwords and keys leaking The kinds of file that hold them are refused before they’re opened, rather than opened and then censored. Anything that does get read is still stripped of secrets on the way out, as a second layer.
Not knowing what happened Every action your AI takes is logged — who, what, when, how long. The log can only be read, never fed back to the AI.
Something going badly wrong Every version of every file is kept, off-site. Updates check their own backup first and undo themselves if they fail.

Found a problem?

Tell us at the address on the contact page. We’ll confirm we got it, and if it’s real it gets tracked the same way everything else here does — which means you can be told exactly what was done about it, and when.