Threat modeling your agent

I was looking at agentic setups recently and ran into awlx’s PicoClaw setup. It’s a local Qwen model on a Mac Studio wired into various personal accounts. It is a well-thought-out build that raises the question I keep getting asked about setups like this: is it safe? The honest answer is “safe against what?”, and that depends far more on who you are than on which model you run. Enter threat modeling, a common practice in the security space that everyone else can learn from!

Threat model

Most agent risk reduces to Simon Willison’s lethal trifecta: access to private data, exposure to untrusted content and a way to send data out. If an agent has all three, anyone who can put text in front of it can ask it to leak whatever it can read. For personal assistants I’d add a fourth leg: actions that cannot be undone. Imagine actions like purchases, sent emails, file deletions and so on. Threat modeling then comes down to four questions: what can the agent reach, who can put text in front of it, how can data leave and what action can’t be undone?

Your laptop is the agent

For anyone running Claude Code or a similar agent in a terminal (especially with --dangerously-skip-permissions), the default answers are uncomfortable.

  • What can the agent reach? The agent runs as you, so it can reach everything your user can: ~/.ssh, ~/.aws, ~/.kube, the gh token, .env files, shell history, browser profiles and every environment variable exported in your shell.
  • Who can put text in front of the agent? Untrusted text arrives through the repositories it reads, the dependencies it installs, web fetches, issue bodies and MCP servers.
  • How can the data leave? curl, git push and any web request are ways out.
  • What action can’t be undone? rm -rf, git push --force, npm publish and a signed transaction are the irreversible actions. The only thing standing between all of that is the permission prompt, which most of us stop reading after the tenth approval (or bypass completely).

How paranoid you need to be therefore depends on the most dangerous thing readable from your home directory. Here’s a quick way to find out:

1
2
ls ~/.ssh ~/.aws ~/.kube ~/.config/gh ~/.docker ~/.npmrc 2>/dev/null
env | grep -iE 'key|token|secret|pass'

If you see something on there you think is sensitive, then it’s worth thinking about how seriously you need to take security for your archetype.

Archetypes

Researchers

By researchers I mean PhD students and anyone else whose work used to live in papers and prototypes but who is now
being dragged into the arena with AI. They typically have few credentials worth stealing, so the realistic risk is
the browser. Session cookies for Google, Overleaf, GitHub, Slack and email skip passwords and 2FA entirely.
Depending on the browser and OS, they sit in the profile directory either in plaintext or one keychain prompt away. Many researchers also publish code and packages that other researchers use, which means you may need to take a peek at the developer threat model even if you don’t identify as one. This archetype might also be enticed into downloading a lot of MCP servers, especially to glue services together, and each of these becomes a new attack vector. Be very, very careful about what you install and regularly audit what you’ve installed.

Developers

Developers have more to lose: SSH keys that push to repositories other people deploy from, signing keys for commits and releases, infrastructure tokens (AWS), package registry tokens and API keys sitting in .env files and exported environment variables. The environment is the easiest leak, since env is a single command and every child process inherits it, so move secrets out of your shell and inject them only into the process that needs them. Put SSH and signing keys behind a password manager that prompts before each use, such as KeePass, 1Password or Bitwarden (extra points for hardware keys). For day-to-day work, Claude Code’s built-in sandbox (/sandbox) is mostly enough; if you’re often swapping agents, then I’d recommend a devcontainer or a VM. There are relatively new solutions like Docker Sandboxes that are worth looking into. I’d also recommend an app like Little Snitch or LuLu to get some local observability.

DevOps/Admins

Admins and DevOps engineers carry crazy high risk because their laptop is a door into everyone else’s systems, and they have privileged access to disable other guardrails. LLM-based debugging has become a default debug pathway for sysadmins, which means log lines, alert payloads and tickets all become injection surfaces. You should definitely be running sandboxed agents at this level, and secrets should be kept out of files by using solutions like Infisical to inject short-lived values per command. Your sandbox should also ideally have an egress protection setup like iron-proxy. SSH and GPG keys should ideally be moved onto hardware keys like a YubiKey, and nothing should be signed without a physical tap.

Give the agent a separate, read-only identity instead of your own, so a leak is limited to what it can see and its actions show up separately in audit logs. By now you should also have a robust CI/CD setup where changes land as PRs, so the agent proposes and a human merges. Finally, keep break-glass access human-only: anything that disables a guardrail should need a step the agent cannot take on its own.

Summary

The framework is short: find the worst thing your agent can read, figure out who can put text in front of it and
understand what can go wrong if each of your secrets is leaked. Try to make it such that the irreversible actions
require human approval (ideally something like a hardware tap) and then let the agents roam free! The same questions apply whether the agent is Claude Code in your terminal or a consumer assistant like Meta’s Muse.