Skip to main content
If your support is already handled by an LLM agent (a chat bot, an email agent, a voice agent), it can use GuidingHand as a tool: when the fix needs hands on the customer’s computer, it sends the invite link, starts a task, relays the agent’s questions to the customer, and reports the result. There are two layers here, and they stay separate:
  • Your agent talks to the customer and decides what should happen.
  • GuidingHand’s agent operates the computer and decides how. Your agent never sees screenshots or clicks; it sends a goal in plain language and gets back progress, questions, approvals and a result.

The loop

  1. Your code (not the model) creates the session with your org API key, gives the model the invite_url to send, and waits for status: connected or the session.connected webhook. Keep the session’s session_token for this conversation.
  2. The model calls start_computer_task with a plain-language prompt, and gets a task id and a cursor.
  3. The model calls wait_for_computer_task with the id and the last cursor. Each call waits up to 25 seconds (timeout_ms, at most 55,000) for new events and returns them.
  4. If the result has a question, the model answers it from what it knows, or asks the customer, and sends question_id and answer in its next wait_for_computer_task call. When question.customer_can_answer is true, the customer sees the question on their screen and may answer it there first; the model can also just keep waiting.
  5. If it has an approval, the model decides (usually by asking the customer) and sends approval_id and decision (approve or deny). When approval.customer_can_approve is true, the customer sees the request on their screen and may allow it or not there first; the model can also just keep waiting.
  6. When done is true, status says how it ended and result or error says what happened. The model tells the customer.

Tool definitions: /tools.json

GET https://guidinghand.ai/tools.json (no authentication) returns ready-made tool definitions for this loop. They describe the session-token task API (/start, /wait, /stop), where the credential is the session_token of one session. That is a good fit for a model: the token can only act on that one computer, never on your org’s other sessions, agents or settings. Each entry has name, method, path, description and, for the first three, an input_schema in JSON Schema. Most LLM APIs take tools in that shape; some call the schema input_schema, others parameters.
Hand each tool call your model makes to runTool and return the JSON to the model as the tool result.

What the tools return

start_computer_task and wait_for_computer_task return the task in the older API’s shape:
events are the events after after; update is the latest one. question or approval is present while the task waits on it; result or error once it ends. Errors are { "error": "message" }, sometimes with more fields such as active_task_id. An answer or decision that comes after the customer already gave one on their screen gets 409 with "code": "already_answered" and "answered_by": "customer": the task has it, so the model just keeps waiting.

Instructions for your model

Tell your model how to use the tools. For example, in its system prompt:
Don’t let your model approve on its own for anything the customer hasn’t agreed to. Approvals exist so a person decides.

Or: your own tools over /v1

If you’d rather define the tools yourself, wrap the /v1 endpoints with your org API key: POST /v1/sessions/{session_id}/tasks to start, GET /v1/tasks/{task_id}/events to follow, POST /v1/tasks/{task_id}/respond to answer, and POST /v1/tasks/{task_id}/stop to stop. Keep the API key in your code and give the model only the arguments it needs (the prompt, the answer), with the session fixed by your code. The org key can reach every session in the org, so it should never be something the model can choose or see.

Tips

  • One task at a time. A computer runs one task at a time. Break big jobs into steps the customer can follow (“First connect to Wi-Fi”, then “Now install the printer driver”).
  • Use request_id. If your agent framework retries tool calls, the same request_id returns the same task instead of starting a second one.
  • Customers can stop it. A task can end stopped because the customer pressed Stop. Treat it as their decision and ask what they’d like to do.
  • Link the replay. Put the task’s replay_url (from GET /v1/tasks/{task_id}) in the ticket so a person can review what happened.