> ## Documentation Index
> Fetch the complete documentation index at: https://docs.guidinghand.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Guardrails

> Rules GuidingHand enforces in code on every step an agent takes: what needs an approval, what never happens, which apps and sites it may use, what it is for, and its limits.

Instructions tell the model how to work. Guardrails are rules GuidingHand enforces in code, on each step the model asks for, before it runs on the computer. The model can't talk its way past them, and neither can anything on the screen. They add to GuidingHand's own rules and can't loosen them.

Every agent starts with these defaults:

```json theme={null}
"guardrails": {
  "mode": "supervised",
  "confirm": ["purchase", "send", "delete", "install", "security", "terms"],
  "block": [],
  "apps": { "allow": [], "block": [] },
  "sites": { "allow": [], "block": [] },
  "safety_checks": "ask",
  "scope": "",
  "screen_check": false,
  "max_steps": 150,
  "max_minutes": null,
  "approval_timeout_minutes": null
}
```

| Key | What it does |
| - | - |
| `mode` | `supervised` (the default): your team, and the customer when `customer_approvals` is on, can approve steps. `unattended`: nobody is there to approve. A step that would need an approval is refused instead, and the agent works around it or finishes and says what's left for a person to do. A safety check stops the task. Questions still go to your team (and the customer, when `customer_answers` is on) as usual. |
| `confirm` | Categories of steps (below) that always wait for an approval. All six by default. |
| `block` | Categories of steps the agent never takes. A category in both lists is blocked; take it out of `block` later and, if it's still in `confirm`, the agent asks first again. |
| `apps.allow`, `apps.block` | App names or ids, such as `Google Chrome`, `com.google.Chrome` or `chrome.exe`, in any case. With an `allow` list, the agent works only in those apps (on a Mac, the Dock, menu bar, Control Center and Spotlight stay usable too: see [below](#apps-and-sites)). Up to 50 each, up to 100 characters. |
| `sites.allow`, `sites.block` | Websites, by host. `example.com` also covers `shop.example.com`. Stored as lowercase hosts: `https://Shop.example.com/cart` is kept as `shop.example.com`. With an `allow` list, the agent acts only on those sites, in a browser or any app showing a web page. Up to 50 each. Site rules need a Mac with GuidingHand 1.0.20 or later for now: see [macOS and Windows](#macos-and-windows). |
| `safety_checks` | When the model raises a safety check on a step: `ask` (the default) for an approval, or `stop` the task. |
| `scope` | What this agent is for, in plain language, up to 2,000 characters. When set, each task is checked against it before it starts, and a task that is clearly something else fails, as does one the check can't give a clear answer on. Empty (the default): no check. |
| `screen_check` | `true`: each time the agent reaches a new app, window or site, GuidingHand checks the screen for text that tries to give the agent instructions (prompt injection). What happens when it finds some, or can't check, follows `safety_checks`. `false` by default. |
| `max_steps` | Model turns per task, 1 to 500. 150 by default. |
| `max_minutes` | Running minutes per task, 1 to 240, not counting time waiting for answers or approvals. `null` (the default): no limit. |
| `approval_timeout_minutes` | An approval nobody decides within this many minutes, 1 to 1,440, is denied. `null` (the default): it waits. |

The categories:

| Category | Steps like |
| - | - |
| `purchase` | Buying or paying: **Buy**, **Pay**, **Place order**, **Subscribe** |
| `send` | Sending, posting or submitting: **Send**, **Post**, **Reply**, **Submit** |
| `delete` | Deleting or removing: **Delete**, **Remove**, **Empty Trash**, **Uninstall** |
| `install` | Installing software: **Install**, **Open Anyway** |
| `security` | Security and account settings: **Sign Out**, **Change Password**, two-factor settings |
| `terms` | Accepting terms or agreements: **I Agree**, **Accept** |

## How each step is checked

Just before a step runs, the computer tells GuidingHand what is there: the app in front, the window title, the page address, what has the focus, and what is under the pointer. GuidingHand then applies these rules in order, and the first one that applies decides:

1. **Secrets, for every agent.** Text that looks like a card number, a US Social Security number, a bank account number (IBAN), an API key, a private key or an access token is not typed, and nothing is typed into a password field. The agent asks the customer to type it themselves.
2. **Blocked apps and sites.** A step in an app on `apps.block`, or on a site on `sites.block`, is refused.
3. **Allow lists.** With a non-empty `apps.allow` or `sites.allow`, a step anywhere else is refused.
4. **Blocked categories.** A step in a `block` category is refused.
5. **Categories to confirm.** A step in a `confirm` category waits for an approval. In `unattended` mode it is refused instead.
6. Anything else runs.

Refusals come first, so a site on both lists is blocked, and so is a category in both lists. Scrolling, moving the pointer and screenshots always run.

### What counts as a step

On a Mac with GuidingHand 1.0.20 or later, GuidingHand judges each step by what it would do:

* **Clicks** by what's under the pointer. A **Place order** button is a `purchase` in any app or site.
* **Keys** by what they act on. Space on a focused button presses it, so it's checked like a click on that button, and so is a space typed while the button has the focus. Typing into a text field isn't. Return or Enter presses the focused button or link. Anywhere else in a Mac dialog, it presses the dialog's default button (such as **Empty Trash**), and is checked as that button. Cmd+Enter or Ctrl+Enter counts as sending, and so does Enter, or a line break in typed text, in a message box (a field labelled like "Message…", "Reply", "Comment" or "Chat"). Delete or Backspace outside a text field, such as on a selected message in Mail, counts as deleting.
* **Drags** at both ends, where they pick up and where they drop. Dropping onto the Trash or Bin counts as deleting, and app rules apply to both ends.
* **Typed text** a line (or field) at a time. It's split at line breaks and tabs, and each piece is checked against the screen as it is then, so a Tab into a password field stops before the password. Text that holds a secret is refused whole: no line of it is typed.
* **Addresses** by where they go. What's typed into a browser's address bar (Chrome, Edge, Safari's smart search field, Firefox) since it was last clicked is judged together, so `evil` then `.test⏎` counts as `evil.test`. With site rules, an address outside the agent's sites is refused before it's typed.
* **Steps GuidingHand can't see.** When the computer can't read what's under a click, or what an acting key (Enter, Space, Delete, Backspace, a typed line break) lands on, and the agent has any `confirm` or `block` categories, a person approves the step first (`Click at (412, 88): GuidingHand couldn’t see what’s there`). In `unattended` mode it's refused.

The model's own approval requests can name a category too: a request in a `block` category is refused without asking anyone, and in `unattended` mode every request is refused.

### Apps and sites

* **Site rules cover anything showing a web page:** every browser, including beta and developer versions, and apps with web content, such as Slack's desktop app, which counts as the site it shows (`app.slack.com`). An app's own local files (`file://`) aren't a site.
* **They fail closed.** With any site rule, allow or block, a step on a page whose address can't be read, or in a browser window that couldn't be read, is refused: GuidingHand can't confirm it isn't a blocked site. Likewise, with any app rule, a step GuidingHand can't place in an app is refused, such as a click on something the computer can't place in one.
* **The system's own apps.** Under an `apps.allow` list, a Mac's Dock, menu bar, Control Center and Spotlight stay usable. On Windows nothing is exempt: the taskbar, File Explorer and the Run box are `explorer.exe`, and Start search is Windows' search process (`SearchHost.exe` or `SearchApp.exe`), so list them to allow them. An `apps.block` entry always wins, even for the Dock.
* **App names differ by system:** `Google Chrome` or `com.google.Chrome` on a Mac, `chrome.exe` on Windows. For computers of both kinds, list both.

## When a step is refused

The model often asks for several steps at once. Each is checked against the screen as it is just before it runs. The steps before a refused one run; the refused step and the ones after it don't, and the model is told exactly which, and why. That doesn't fail the task: the agent plans around it, or finishes and says what's left for a person to do. The agent's instructions list its guardrails, so it seldom tries.

A step held for approval is an ordinary approval, not a `blocked` event. Its `approval_required` event has `source: "guardrail"` and `risk: "high"`, and its `action` is GuidingHand's own description of the step, not the model's, for example `Click “Place order” in Safari (shop.example.com)`. If the agent asked for approval itself just before and named the same category (it asked about `send`, then clicks **Send**), that approval covers GuidingHand's check for that one step. Otherwise GuidingHand asks again, in its own words.

Every refusal or stop is a [`blocked` event](/concepts/tasks#events) on the task. `message` says what didn't happen and why, and `data.guardrail` says which rule decided:

```json theme={null}
{ "cursor": 12, "type": "blocked", "message": "Didn’t click “Place order”: this agent can’t make purchases or payments.", "ts": "2026-09-29T15:05:39.584Z",
  "data": { "guardrail": { "rule": "category", "decision": "block", "category": "purchase", "app": "Safari", "site": "shop.example.com" } } }
```

| `decision` | What happened |
| - | - |
| `block` | The step didn't run. The agent carries on without it. |
| `handoff` | A secret, or a password field: nothing was typed. The agent asks the customer to type it themselves. |
| `refuse` | A request was refused without asking anyone: an approval this agent can't get, a task outside its `scope`, or a task on a computer that can't check the agent's rules. |
| `stop` | The task ended: a safety check, a screen check or a limit stopped it, or (on an OpenAI model) OpenAI's safety monitor did. |

`rule` is `secret`, `password_field`, `apps`, `sites`, `category`, `unattended`, `context`, `scope`, `limit`, `safety_check`, `screen_check` or `openai_monitor`. `category`, `app` and `site` are there when they apply. A secret itself is never recorded. The task's trace and GuidingHand's server log name only the rule ("Stopped by the agent’s guardrails (scope)."): the reasons, and any text quoted from the screen, are in the task's `error` and events.

On the customer's screen, with `narration` on, the app says what the agent didn't do and why, in the words of the event's `message`. Your team sees the same in the console, in the task's timeline and replay.

## Timeouts, limits and stops

* **Approval timeout.** An approval nobody decides within `approval_timeout_minutes` is denied. The `denied` event has `answered_by: "timeout"`, and its message says how long it waited ("Denied: nobody decided within 10 minutes."). The [`task.approval_decided`](/guides/webhooks#event-types) webhook has `answered_by: "timeout"` and the same in its `note`. A decision sent after that gets `409` with `code: "already_answered"`, `answered_by: "timeout"` and the message "That approval timed out: nobody decided in time, so it was denied." A safety check that times out is denied like any other, which stops the task.
* **Limits.** A task that reaches `max_steps` or `max_minutes` fails, and `error` says which limit it reached. Break long jobs into several tasks.
* **Scope.** A task that is clearly outside `scope` fails before the agent takes a step, and `error` says why. So does a task the check can't give a clear yes or no on. It costs nothing: the agent never acted.
* **Safety checks.** With `safety_checks: "stop"`, or in `unattended` mode, a safety check ends the task `stopped`, with the reason in `error`.
* **Screen check.** When `screen_check` finds text aimed at the agent, or can't check a screen, it asks for an approval (`safety_checks: "ask"`) or ends the task `stopped` (`"stop"`, or `unattended` mode). Denying that approval stops the task too.
* **The model's own safety systems.** On an OpenAI model, OpenAI's safety monitor can end a conversation it judges unsafe (`rule: "openai_monitor"`). A Claude model can decline to go on. Either way the task fails with the reason in `error`, and can't be resumed.

## macOS and Windows

What the computer can tell GuidingHand depends on its system, with GuidingHand 1.0.20 or later:

| | macOS | Windows |
| - | - | - |
| The app in front and its window title | Yes | Yes |
| The page address in a browser | Yes | Not yet |
| What's under the pointer and what has the focus, including password fields | Yes | Not yet |

So on Windows, for now:

* **App rules work.** A click counts as in the app in front.
* **Site rules don't.** A task with an agent that has site rules fails at the start, with the reason in `error`, rather than running unchecked. Use an agent without site rules for Windows computers.
* **Categories rest on the agent.** GuidingHand can't see what's under a click or what has the focus, so `confirm` and `block` rely on the agent's own approval requests, which name a category, and on its instructions. Secrets are still caught in typed text, but a password field isn't recognized.

Before 1.0.20, the GuidingHand app reports none of this, on either system: a task with an agent that has app or site rules fails at the start, with an `error` that asks the customer to update GuidingHand, and categories and password fields rest on the agent, as on Windows.

On a Mac, GuidingHand reads the screen's controls through Accessibility. If the customer hasn't allowed it, GuidingHand can't check the agent's guardrails, so a task with rules to check fails at the start (a `blocked` event with `rule: "context"` and `decision: "refuse"`, and an `error` that says how to turn it on) rather than asking about every click.

Known limits:

* The computer reads the address of a browser's front window, so a click in a second window of the same browser is judged by the front window's address.
* Enter in an ordinary form field submits the form, and isn't treated as sending: only message boxes are. The agent's own approval requests cover form submissions.

## Changing guardrails

`PATCH /v1/agents/{agent_id}` takes a partial `guardrails`: only the keys you send change. Inside `apps` and `sites`, only the list you send changes. A key set to `null` goes back to its default, and `"guardrails": null` resets them all. `POST /v1/agents` takes the same. A value out of range, an unknown key or a site that isn't a host address is a `400` that says which.

```bash theme={null}
curl -X PATCH https://guidinghand.ai/v1/agents/billing \
  -H "Authorization: Bearer $GUIDINGHAND_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "guardrails": {
      "block": ["purchase", "install"],
      "apps": { "allow": ["Acme Billing", "Safari"] },
      "sites": { "allow": ["acme.com"] },
      "scope": "Billing help in Acme Billing and on acme.com: invoices, billing addresses and plan details.",
      "approval_timeout_minutes": 10
    }
  }'
```

The response is the whole agent, with every guardrail filled in. Lists come back as you set them: after the example above, `confirm` still has all six categories, and `purchase` and `install` are blocked because `block` wins. You can also set guardrails in the console under **Agents**. Like the other settings, they are read when a task starts: a change applies to the next task, not one already running. The org's [audit log](/concepts/orgs#audit-log) records each change, with the guardrail's value before and after.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.