Skip to content

Agent security

Donkit Guard watches two moments of every agent turn: what the agent is about to send to a model, and what it is about to do with a tool. Depending on the mode you choose for your account, it records what it sees, masks secrets and personal data, and asks you before an agent takes an action with consequences.

What it does

  • Keeps secrets out of models. API keys, tokens, passwords and similar credentials found in a message, a document or a tool result are replaced with a placeholder before the text reaches any model. The original text stays in your conversation; only the copy sent to the model is masked.
  • Keeps personal data inside the platform. Names, contact details and identification numbers are masked before they reach a model that runs outside the platform, such as your own model provider. Platform models receive them unchanged, and the finding is recorded.
  • Asks before side effects. Deleting or editing files, sending email, writing to a spreadsheet, running code with network access, spending credits on media — the agent pauses and shows an approval card in the chat to the person who asked for it. They Allow or Reject; if nobody answers within five minutes, the action is not run. Reading files, searching, and running code in the sandbox without network access do not ask.
  • Records every decision. The Security page lists what the guard allowed, masked, blocked or asked about — with the agent, the user, the tool or model, the rule and the reason. You can narrow the list by outcome.
  • Follows your rules. An account owner or admin describes the organization's own rules in plain words — forbid a tool, make an action ask for confirmation, mask more than the platform does — tries them on one action, and every agent of the account follows them.

Modes

Open Settings → Security. The mode applies to every agent of the account within about half a minute:

  • Off — nothing is checked or recorded.
  • Observe — everything is checked and recorded, nothing is blocked or masked. Use it first to see what the guard would do; the list marks the entries that Enforce would have blocked.
  • Enforce — secrets are masked, personal data is masked before outside models, and actions with side effects wait for your approval.

New accounts start in the platform default, which the Security page shows: Off, unless your installation sets another one. Only an account owner or admin can change the mode. Everything on this page applies to an account from the moment it is in Observe or Enforce.

Your own rules

Below the mode, the Security page asks you to describe what your agents may and may not do — in your own words, one thought per line:

Nobody may delete files.
Sending email always asks the person first.
No customer data to models outside the platform.

Draft rules turns the description into a list of rules and shows each one before anything is saved. Every card says, in one sentence, what the rule does — Blocks the tools delete_file. Nobody can confirm past it. — along with the reason a user will read when the rule stops them. Keep the ones you want with the checkbox and leave the rest out.

Two things a card can say instead:

  • This one cannot be used. The rule named a tool your account does not have, reused a platform rule's name, or left out the reason a blocking rule needs. The problems are listed and the rule cannot be kept until the description says something the guard can actually do.
  • No rule can express this. Your sentence is quoted back with why it does not fit — the guard sees roles and tools, not job titles or people by name — and the nearest rules it can express are offered to pick from, or to skip. Nothing is silently substituted for what you asked.

Your rules are combined with the platform rules, and the stricter setting always wins: a rule of yours can forbid an action or make it ask for confirmation, but it can never relax a platform rule (secrets stay masked, side effects still ask). The mode is set with the switch above, not by a rule.

Every organization starts with one rule of its own: people outside it — visitors of a published agent, API clients and other agents calling yours — may not take actions with side effects. It shows up among your rules like any other, and you can delete it; visitors then confirm such actions themselves.

Drafting asks a model to read your description, so the wording of the answer varies and the hourly number of drafts per account is limited. What the page shows is not the model's word for it: the sentence on each card is built from the rule itself, and the platform checks every rule before you can keep it.

Trying a rule before you save it

Try it before saving answers what the guard would do with one concrete action under the rules you have just assembled — before any agent runs under them. Pick a tool and its arguments, or write a message an agent would send to a model, and press Run. The answer is Blocked, Asks the person, Masked or Allowed, with the rule that decided it and the reason. A masked message also shows the text the model would have received. Nothing is sent to a model and nothing is added to the decision list.

While the account is in Observe or Off, the dry run answers as Enforce would — it is there to tell you what switching the mode on will change.

The rules you have

The Organization rules card lists the rules the account runs on now, one card each, with Delete to drop one. While you have unsaved changes the card shows the rules the account would run on, and Save as version N stores them; Discard throws the changes away. A save applies to every agent of the account within about half a minute, and cancels approvals still waiting for an answer — they were asked under the previous rules.

Advanced: the document itself

Rules are stored as a short YAML document, and Advanced at the bottom of the page opens it: the text editor, the version history and the full table of effective rules. Edit it directly when a rule needs a selector the description cannot reach.

Each rule has an id (unique, lowercase, dashes), a match block, an action and a reason. deny with hard: true cannot be overridden by an approval. Three examples:

schema_version: 1
layers:
  - name: organization
    rules:
      # Agents never write to Jira; the glob matches every tool of the server.
      - id: org-no-jira-writes
        match: {action_kind: [tool_call], tool: ["jira_*"]}
        action: deny
        hard: true
        reason: "Jira changes are made by people, not agents"

      # Running code always asks, even without network access.
      - id: org-python-needs-approval
        match: {action_kind: [tool_call], tool: [python_exec, run_file]}
        action: require_approval
        reason: "Running code needs confirmation in this organization"

      # Personal data is masked before every model, platform models included.
      - id: org-mask-pii-everywhere
        match:
          action_kind: [llm_request]
          finding_class: [pii]
          min_score: 0.5
          destination_trust: [platform, tenant, unknown]
        action: mask
        reason: "Personal data is masked before any model"

The selectors you can use in match are listed in the document's own comments (tool name, effect, who is acting, where the data goes, what the detectors found). Check validates the text and previews the rules it adds; Save stores it as a new version. Every problem is reported with where it sits — the line for a syntax error, the path of the rule (layers/0/rules/1/reason) for anything else — and a document that does not validate is never applied.

The version history keeps every saved document; Roll back makes an earlier one active again. The table of effective rules shows all rules that apply to the account — the platform's, which you cannot edit, and yours — each as a sentence with the raw selectors beneath it. When several rules match one action, the strictest one decides.

Who can approve

An approval confirms what you asked the agent to do, so the card goes to the person whose message the agent is working on, and only they can answer it:

  • You and the members of your organization confirm your own actions, in the chat where the agent asked, and in the Builder while you are building. A member of the organization using its agent gets the card, not the agent's owner.
  • Owners and admins of the organization do not approve for anyone else. Their say over what agents may do is the rules on the Security page: a rule can forbid an action outright or make it ask, and nobody can confirm their way past a rule that forbids.
  • Visitors of a published agent — people outside your organization — cannot take actions with side effects at all while the rule the organization starts with is in place (see Your own rules). Delete that rule, and a visitor confirms such actions themselves, the same way you do.
  • API clients and other agents calling yours have no one to ask: in Enforce mode an action of theirs that would need approval is blocked.
  • Workflow runs have no one to ask either: in Enforce mode a workflow step's action that would need approval — a write through Google Workspace or an MCP server, for example — is blocked. The run stops under Step refused with This step was refused — by a workflow policy or by the service it called. and waits for you. The step's error gives the guard's reason and says that a workflow run can't ask for approval yet: to let the step run, allow the action in your rules on the Security page, or do it yourself outside the workflow. The run does not retry or try to fix itself.

What you see when something is blocked

  • In the chat, a notice explains that the guard blocked the action, that the approval was rejected, or that no decision came in time. The agent's turn ends there.
  • In some cases the agent is told the action was blocked and continues, explaining what it could not do. The blocked tool call shows Blocked with the reason in place of its normal result.
  • In Enforce mode, a request the guard could not finish checking is not let through unchecked. The notice says so and asks you to try again: no rule of yours stopped it, and the same request usually goes through once the check completes. The guard waits for a check as long as it takes; it gives up only when the checking service stops answering altogether.
  • In a workflow run, when the guard blocks what a step is doing — the action of an Action step, or a request a step sends to a model — the run stops under Step refused and waits for you, and the step's error gives the guard's reason: a rule of yours that forbids it, or an approval the run could not ask for. When the guard could not finish checking it, the step is retried automatically instead, like any temporary failure, and a scheduled workflow keeps its schedule; only if the check still cannot finish after the automatic retries does the run stop, under Retries ran out. The tools an Agent step calls are handled as in the chat: a tool call the guard refuses in a way that ends the turn stops the run under Step refused, and one it blocks without ending the turn, or could not finish checking, is reported to the agent, which carries on.

Mobile apps

The approval card is available in the web app and the Builder. The mobile apps show the request as a plain question for now.

Agent security is off until your account switches it on: the platform default is Off unless your installation sets another one. An account in Observe or Enforce gets everything described on this page, workflow runs included.