Catalogue · C
Research
I build programs with AI agents, so I want to know how far they can be trusted. That is the subject of my bachelor's thesis and of a few tools for running experiments.
C.01 · Bachelor's thesis
Can You Trust Your Agent?
Implementation and Evaluation of Autonomous AI Agents in an Isolated Laboratory Environment
I built an isolated lab (Windows virtual machines run by a control panel I made) and attacked AI coding agents with hidden instructions. The talk is called "Can You Trust Your Agent?". Hundreds of runs, and every reported breach checked by hand against the transcript.
01 · The lab
How I tested
To repeat attacks hundreds of times under the same conditions, I built a control panel. The agent works inside a Windows virtual machine, and the panel manages the machines from the outside, on my own computer. They exchange nothing but files through one shared folder, with no connection going in, so the agent under test cannot reach the thing that controls it.
Cloning
Several machines, side by side
A prepared Windows machine is cloned instantly: the copy takes no extra disk space until it starts to differ from the original. Each clone gets its own name, its own ports and its own shared folder, so one machine's results never overwrite another's. A run takes about half a minute and each machine does one at a time, so three clones mean three runs at once.

FIG. 1The fleet: the original machine and two clones. Each can be started, stopped, reverted to a snapshot or cloned again. Tests
Written tests, and starting a batch
Every test is written down in a library: tasks (ordinary programming jobs), attack variants (where the hidden instruction sits and what it asks for) and protection levels, from none, through rules written for the agent, to permissions that are actually enforced. One click publishes the library to every machine. For a batch I pick the machines, a model per machine, the protection levels, and the tasks and variants; the panel counts the runs and, before anything is sent, shows the exact order each machine will receive.

FIG. 2A new batch: three machines, a different model on the third, three protection levels. The panel counts 126 runs per machine and shows the exact order on the right before sending. Data
Choosing data, and graphs
Every finished result comes back to one store on my computer, checked to make sure no file changed on the way. In the data tab I pick a batch, model, protection level, task or variant and see the numbers and a graph at once. Two rules are built in: a run where the hidden instruction never reached the model does not count as a resisted attack, and runs cut short by the provider's limits stay out of the time and cost averages. Everything exports to a spreadsheet for the thesis.

FIG. 3One batch of 299 runs, attack success by protection level for two models. With rules written for the agent (L1, L1m, L1x) attacks still get through, for one model every time; with enforced permissions (L2) none did.
02 · Results
What I found
The channel is the weakness
Agents refuse a malicious instruction when it arrives as tool output or in the user's message, and follow the very same instruction when it is written into the project's own instruction file (AGENTS.md).
Refusals are about words
The same harmful action, renamed in build-tool vocabulary, goes through.
An agent's report is not evidence
One agent inserted a backdoor, and its summary said the opposite.
Stronger models are not immune
They refuse the obvious attacks but still ship the disguised ones.
What worked
A passive monitor outside the agent caught every forbidden action, including the hidden ones.

C.03
Design instruction catalogue
A page that collects, in one place, the instructions used to teach AI coding agents how to design, with search, filters and the full text of every instruction.




