Everything on this page is something you can ask to see.
The sessions below come from our own instance, which runs more than ships today — what ships is on the front page.
The same question, asked twice
One question, put to the same assistant twice on 9 August 2026. On the left it has no
record and says so — it ends by asking which amplifier the building has. On the right
it has the vault, and opens with the amplifier by make and model, the extension, the exact
port the line was moved to in June, and the repair history that explains why.
Nothing changed but the record it was allowed to read. Every session
below ran on Grok; the June repair notes it is quoting were filed months earlier by a
different assistant entirely.


The client is anonymised and one purchase-order number is blacked out. Both screens were
rebuilt at readable size from the original session — the words are the assistant’s, and
every fact in the right-hand answer is filed somewhere we can show you.
What you get back when you finish
You type one word. A receipt comes back, shaped like this one:
Saved. The full conversation, kept word for word
Filed. Sorted into the right customer's history
Checked. 99.7% of it accounted for — nothing quietly dropped
Closed. The job you were working on, marked done
That is one real receipt, not a published average — the figure on the third line is measured for your session, every time, and it is the only number on this page for exactly that reason.
The third line is the one that matters. A summary that loses half your work but reports success is worse than no summary at all, because you won’t find out until you need it. So VaultHelm measures how much of what you gave it actually made it through. If too much went missing, it keeps everything word for word instead. It is not permitted to report success without doing that check.
And if your connection drops mid-save, just say the word again. A repeat only fills in what didn’t make it — it can’t create a duplicate and it can’t overwrite what’s already there, and that is a property of the thing rather than a habit we asked it to keep.
Finding something again is three steps, not a guess
When you ask what was done for a customer, you don’t get an AI’s best recollection. You get:
- The entry — who, when, what it was about.
- The write-up — what was actually done, in that customer’s history.
- The original — the raw conversation, exactly as it happened.
You can always reach the third one. The middle one can be corrected if it got something wrong. The one underneath is never rewritten — which is the difference between a record you can rely on and a summary you have to trust.
Open the folder yourself
On the walkthrough we do this in front of you, because it is the part that settles the question.
Split the screen. On one side, an assistant quoting a paragraph back at you. On the other, that same paragraph sitting in a text file on a laptop, in a folder you could copy to a memory stick — no product running, nothing connected, nothing logged in.
Both halves are the same words, because the file is where the words were. The folder is the product. The chat is only a window onto it. Every claim on this page about ownership and lock-in comes down to that one demonstration, and it takes about a minute.
“What’s still open, and did anyone go back?”
The headline comparison shows recall. This is the one that shows judgement, and it is the harder trick.
Take an incident from a fortnight earlier and the loose ends it left. Ask what happened to them. The answer goes item by item against everything filed since — and for several of them the answer is that nothing was ever filed, so nobody went back.
That reply does not flatter the shop that asked for it, which is the point. An assistant guessing from a chat window has every incentive to produce a tidy summary where the loose ends look handled. Silence has to be a first-class answer, or the record is worse than useless — it’s reassuring.
More from the same session
Smaller than the headline comparison, and in some ways more telling — this is what the discipline looks like on an ordinary evening.
It withdrew its own alarming finding
It reported that a set of machines was down. Told that didn’t sound right, it didn’t argue and it didn’t fold politely either — it went and checked each one, then retracted its own headline in writing: the claim was false, here is what is actually true, and here is the source I should have read instead of the one I did. The wrong version stays in the record next to the correction.
It misheard the company name. Three times.
The whole session was dictated, and three times a company or a person’s name came through mangled — close enough to sound real, wrong enough to match nothing. Each time it found the name the garble was reaching for, said plainly which one it had landed on, and answered about that. It never invented a customer to fit the question.
It hit a wall and said so
Asked about something in a file it isn’t allowed to open, it doesn’t guess and it doesn’t quietly answer from something adjacent. It tells you it couldn’t look, and which door was shut. A refusal you can see beats a confident answer built on a source that was never read.
“Now write it up”
Scattered notes from half a dozen visits became a full phased deployment plan — what’s done, what’s blocked, what has to happen in what order, and the decisions only a person standing on site can make. Then it filed itself in the right client folder.
What we deliberately don’t measure
You won’t find uptime percentages, customer counts, speed benchmarks or named references on this page.
Uptime and speed figures would need measurement we haven’t built to a standard worth publishing, and a number we can’t stand behind is worse than no number. Customer detail isn’t ours to put on a website.
If something matters to your decision and it isn’t here, ask. If it doesn’t exist yet, that’s the answer you’ll get.