> ## Documentation Index
> Fetch the complete documentation index at: https://docs.guidinghand.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Drive GuidingHand from your AI agent

> Give your own LLM support agent tools to start, follow and stop tasks on the customer's computer.

If your support is already handled by an LLM agent (a chat bot, an email agent, a voice agent), it can use GuidingHand as a tool: when the fix needs hands on the customer's computer, it sends the invite link, starts a task, relays the agent's questions to the customer, and reports the result.

There are two layers here, and they stay separate:

* **Your agent** talks to the customer and decides *what* should happen.
* **GuidingHand's agent** operates the computer and decides *how*. Your agent never sees screenshots or clicks; it sends a goal in plain language and gets back progress, questions, approvals and a result.

## The loop

```mermaid theme={null}
sequenceDiagram
    participant C as Customer
    participant Y as Your agent
    participant G as GuidingHand

    Y->>G: POST /v1/sessions (your code, org API key)
    G-->>Y: invite_url, session_token
    Y->>C: "Open this link and click Connect"
    G-->>Y: session.connected
    Y->>G: start_computer_task {prompt}
    loop until done
        Y->>G: wait_for_computer_task {id, after}
        G-->>Y: events, status, question or approval
        opt question
            Y->>C: relays the question (or answers from context)
            Y->>G: wait_for_computer_task {id, after, question_id, answer}
        end
    end
    Y->>C: tells the customer the result
```

1. **Your code** (not the model) creates the session with your org API key, gives the model the `invite_url` to send, and waits for `status: connected` or the `session.connected` webhook. Keep the session's `session_token` for this conversation.
2. The model calls `start_computer_task` with a plain-language `prompt`, and gets a task `id` and a `cursor`.
3. The model calls `wait_for_computer_task` with the `id` and the last `cursor`. Each call waits up to 25 seconds (`timeout_ms`, at most 55,000) for new events and returns them.
4. If the result has a `question`, the model answers it from what it knows, or asks the customer, and sends `question_id` and `answer` in its next `wait_for_computer_task` call. When `question.customer_can_answer` is `true`, the customer sees the question on their screen and may answer it there first; the model can also just keep waiting.
5. If it has an `approval`, the model decides (usually by asking the customer) and sends `approval_id` and `decision` (`approve` or `deny`). When `approval.customer_can_approve` is `true`, the customer sees the request on their screen and may allow it or not there first; the model can also just keep waiting.
6. When `done` is `true`, `status` says how it ended and `result` or `error` says what happened. The model tells the customer.

## Tool definitions: `/tools.json`

`GET https://guidinghand.ai/tools.json` (no authentication) returns ready-made tool definitions for this loop. They describe the session-token task API (`/start`, `/wait`, `/stop`), where the credential is the `session_token` of one session. That is a good fit for a model: the token can only act on that one computer, never on your org's other sessions, agents or settings.

| Tool                     | Endpoint      | Input                                                                                                                |
| ------------------------ | ------------- | -------------------------------------------------------------------------------------------------------------------- |
| `start_computer_task`    | `POST /start` | `prompt` (required), `request_id`                                                                                    |
| `wait_for_computer_task` | `POST /wait`  | `id` (required), `after`, `timeout_ms`, and either `question_id` + `answer` or `approval_id` + `decision` (+ `note`) |
| `stop_computer_task`     | `POST /stop`  | `id` (required)                                                                                                      |
| `health`                 | `GET /health` | none. With the token, also reports whether the computer is connected.                                                |

Each entry has `name`, `method`, `path`, `description` and, for the first three, an `input_schema` in JSON Schema. Most LLM APIs take tools in that shape; some call the schema `input_schema`, others `parameters`.

```javascript theme={null}
const BASE = 'https://guidinghand.ai';
const manifest = await (await fetch(`${BASE}/tools.json`)).json();

// Tool definitions to pass to your model.
const tools = manifest.endpoints
  .filter((e) => e.input_schema)
  .map((e) => ({ name: e.name, description: e.description, input_schema: e.input_schema }));
const byName = Object.fromEntries(manifest.endpoints.map((e) => [e.name, e]));

// Runs one tool call from the model, on this conversation's session.
async function runTool(name, input, sessionToken) {
  const e = byName[name];
  const res = await fetch(BASE + e.path, {
    method: e.method,
    headers: { Authorization: `Bearer ${sessionToken}`, 'Content-Type': 'application/json' },
    body: e.method === 'POST' ? JSON.stringify(input) : undefined,
  });
  return res.json(); // errors too: the model can read { "error": "..." } and adjust
}
```

Hand each tool call your model makes to `runTool` and return the JSON to the model as the tool result.

### What the tools return

`start_computer_task` and `wait_for_computer_task` return the task in the older API's shape:

```json theme={null}
{
  "id": "task_RMW0Ke37PoUW",
  "request_id": "T-1042-dark-mode",
  "status": "waiting_for_user",
  "done": false,
  "cursor": 4,
  "events": [
    { "cursor": 4, "type": "question", "message": "Which account?", "ts": "2026-09-26T15:05:41.593Z",
      "data": { "question_id": "q_nh8TwFRVQMs", "question": "Which account?", "options": ["Work", "Home"] } }
  ],
  "update": { "cursor": 4, "type": "question", "message": "Which account?", "ts": "2026-09-26T15:05:41.593Z" },
  "question": { "question_id": "q_nh8TwFRVQMs", "question": "Which account?", "options": ["Work", "Home"], "customer_can_answer": true }
}
```

`events` are the events after `after`; `update` is the latest one. `question` or `approval` is present while the task waits on it; `result` or `error` once it ends. Errors are `{ "error": "message" }`, sometimes with more fields such as `active_task_id`. An answer or decision that comes after the customer already gave one on their screen gets `409` with `"code": "already_answered"` and `"answered_by": "customer"`: the task has it, so the model just keeps waiting.

## Instructions for your model

Tell your model how to use the tools. For example, in its system prompt:

```text theme={null}
You can operate the customer's computer with the computer-task tools, once they have connected.
- Describe the outcome in start_computer_task, not the clicks: "Turn on Dark Mode", not "click the Apple menu".
- Keep calling wait_for_computer_task with the latest cursor until done is true.
- When the task has a question, answer it if the conversation already tells you the answer. Otherwise ask the customer, in their words, and pass their answer on. If question.customer_can_answer is true, they can also answer it on their screen: tell them so, and keep waiting.
- If an answer or decision comes back with code "already_answered" or "not_pending", the customer answered first (or it already closed). That's fine: keep waiting.
- When the task asks for approval, tell the customer exactly what will happen and approve only if they agree. If approval.customer_can_approve is true, they can also allow it on their screen: tell them so, and keep waiting.
- Never ask the customer for passwords, codes or card numbers to pass on. The computer agent never types them or asks for them; when a step needs one, its question asks the customer to type it on their own computer. Relay that, then answer the question once they have.
- If the customer wants it to stop, call stop_computer_task.
```

Don't let your model approve on its own for anything the customer hasn't agreed to. Approvals exist so a person decides.

## Or: your own tools over `/v1`

If you'd rather define the tools yourself, wrap the `/v1` endpoints with your org API key: `POST /v1/sessions/{session_id}/tasks` to start, `GET /v1/tasks/{task_id}/events` to follow, `POST /v1/tasks/{task_id}/respond` to answer, and `POST /v1/tasks/{task_id}/stop` to stop. Keep the API key in your code and give the model only the arguments it needs (the prompt, the answer), with the session fixed by your code. The org key can reach every session in the org, so it should never be something the model can choose or see.

## Tips

* **One task at a time.** A computer runs one task at a time. Break big jobs into steps the customer can follow ("First connect to Wi-Fi", then "Now install the printer driver").
* **Use `request_id`.** If your agent framework retries tool calls, the same `request_id` returns the same task instead of starting a second one.
* **Customers can stop it.** A task can end `stopped` because the customer pressed Stop. Treat it as their decision and ask what they'd like to do.
* **Link the replay.** Put the task's `replay_url` (from `GET /v1/tasks/{task_id}`) in the ticket so a person can review what happened.
