- Your agent talks to the customer and decides what should happen.
- GuidingHand’s agent operates the computer and decides how. Your agent never sees screenshots or clicks; it sends a goal in plain language and gets back progress, questions, approvals and a result.
The loop
- Your code (not the model) creates the session with your org API key, gives the model the
invite_urlto send, and waits forstatus: connectedor thesession.connectedwebhook. Keep the session’ssession_tokenfor this conversation. - The model calls
start_computer_taskwith a plain-languageprompt, and gets a taskidand acursor. - The model calls
wait_for_computer_taskwith theidand the lastcursor. Each call waits up to 25 seconds (timeout_ms, at most 55,000) for new events and returns them. - If the result has a
question, the model answers it from what it knows, or asks the customer, and sendsquestion_idandanswerin its nextwait_for_computer_taskcall. Whenquestion.customer_can_answeristrue, the customer sees the question on their screen and may answer it there first; the model can also just keep waiting. - If it has an
approval, the model decides (usually by asking the customer) and sendsapproval_idanddecision(approveordeny). Whenapproval.customer_can_approveistrue, the customer sees the request on their screen and may allow it or not there first; the model can also just keep waiting. - When
doneistrue,statussays how it ended andresultorerrorsays what happened. The model tells the customer.
Tool definitions: /tools.json
GET https://guidinghand.ai/tools.json (no authentication) returns ready-made tool definitions for this loop. They describe the session-token task API (/start, /wait, /stop), where the credential is the session_token of one session. That is a good fit for a model: the token can only act on that one computer, never on your org’s other sessions, agents or settings.
Each entry has
name, method, path, description and, for the first three, an input_schema in JSON Schema. Most LLM APIs take tools in that shape; some call the schema input_schema, others parameters.
runTool and return the JSON to the model as the tool result.
What the tools return
start_computer_task and wait_for_computer_task return the task in the older API’s shape:
events are the events after after; update is the latest one. question or approval is present while the task waits on it; result or error once it ends. Errors are { "error": "message" }, sometimes with more fields such as active_task_id. An answer or decision that comes after the customer already gave one on their screen gets 409 with "code": "already_answered" and "answered_by": "customer": the task has it, so the model just keeps waiting.
Instructions for your model
Tell your model how to use the tools. For example, in its system prompt:Or: your own tools over /v1
If you’d rather define the tools yourself, wrap the /v1 endpoints with your org API key: POST /v1/sessions/{session_id}/tasks to start, GET /v1/tasks/{task_id}/events to follow, POST /v1/tasks/{task_id}/respond to answer, and POST /v1/tasks/{task_id}/stop to stop. Keep the API key in your code and give the model only the arguments it needs (the prompt, the answer), with the session fixed by your code. The org key can reach every session in the org, so it should never be something the model can choose or see.
Tips
- One task at a time. A computer runs one task at a time. Break big jobs into steps the customer can follow (“First connect to Wi-Fi”, then “Now install the printer driver”).
- Use
request_id. If your agent framework retries tool calls, the samerequest_idreturns the same task instead of starting a second one. - Customers can stop it. A task can end
stoppedbecause the customer pressed Stop. Treat it as their decision and ask what they’d like to do. - Link the replay. Put the task’s
replay_url(fromGET /v1/tasks/{task_id}) in the ticket so a person can review what happened.