/v1 API, both SDKs and the console’s Evals tab.
Test cases
A test case is a problem on a machine and a way to tell it’s fixed:
Create cases with
POST /v1/test_cases, change them with PATCH /v1/test_cases/{test_case_id} (only the fields you send change), list them with GET /v1/test_cases?tag=&category=&test_set_id=&q=, and delete them with DELETE /v1/test_cases/{test_case_id}.
To load a whole benchmark, POST /v1/test_cases/batch creates or updates up to 500 cases at once, matched by key, and with test_set_id adds them to that set. Run it again after editing your file: cases with a known key are updated, new ones are created. It returns { "data", "created", "updated" }.
How a try is scored
The machine’s end state decides. The agent saying “Done.” never counts.- A check script runs on the machine (elevated, as the signed-in user) and prints
Trueonce the problem is fixed. It runs twice: before the task, where it must printFalse(proof the setup really broke the machine), and after it. For a case that asks a question rather than fixes something, the script reads the agent’s final answer from theGH_AGENT_ANSWERenvironment variable.GH_CHECK_PHASEsays which run it is (beforeorafter): a case where nothing is broken beforehand, like a risky request the agent should refuse, printsFalsewhen it’sbefore. - A rubric says what a fix looks like in plain language. A model judges it from the recorded conversation and the first and last screens, and gives a
reason. - A case can have both. An attempt passes only when each one it has passes.
outcome:
The simulated customer
During a test, nobody sits at the machine. A model plays the customer: it gets the case’spersona, its facts, the agent’s question and the machine’s screen, and answers the way that person would. It never sees the expected fix or the check.
factsare what the customer knows, by name (up to 50). A value written as{ "secret": "…" }is never shown to the model: it only knows the fact’s name, and can type the value into the field that has focus, the way a person types a password.clicks_admin_prompts: truelets it click Yes on an admin (UAC) prompt when the agent asks.does_steps: falsemakes it hand steps back (“Could you do that for me?”) when the agent asks it to do the fix itself.
operator.approvals. Simulated answers are marked simulated in the task’s events, and each attempt lists the customer’s turns in customer.
Test sets
A test set is a list of cases to run together, with a default number of tries (repeat, 1 to 20) and a default customer for its cases that have none.
POST /v1/test_sets/{test_set_id}/cases adds cases (or changes the repeat of ones already in it), DELETE /v1/test_sets/{test_set_id}/cases/{test_case_id} takes one out, and PATCH with cases replaces the whole list. Deleting a set keeps its cases and its runs.
Runs and attempts
POST /v1/test_runs runs a test set (test_set_id) or a list of cases (test_case_ids) against each of agents (default ["default"]). Tries per case are the first of:
repeat_by_case[case]in the request,repeatin the request,- the case’s
repeatin the set, - the set’s
repeat, -
queued, then running, then completed once none of its attempts are left (or cancelled). test_run.completed is sent to your webhook when it ends.
GET /v1/test_runs/{test_run_id}/resultshas a row per case and agent (tries,passed,failed,error,all_passed,median_seconds,cost_cents) and breakdownsby_agent,by_tagandby_category. It counts the last attempt of each try, so an error that was retried counts as its retry.GET /v1/test_runs/{test_run_id}/attempts?test_case_id=&agent_id=&outcome=lists the attempts. Each has the check’s output before and after (check), the judge’s verdict and reason (judge), the simulated customer’s turns (customer), itsmetrics(active seconds, steps, questions, approvals, cost) and thereplay_urlof its task.GET /v1/attempts/{attempt_id}is one attempt, with the copy of the test case it ran.POST /v1/test_runs/{test_run_id}/cancelstops a run.POST /v1/test_runs/{test_run_id}/rerunwith{ "only": "failed" }(or"error", or"all") starts a new run of those (case, agent) pairs.
Quick start
Create a case, put it in a set, run the set against two agents, wait, and read the results.1
Create a test case
2
Put it in a test set
3
Run it against your agents
Three tries per case for each of two agents: six attempts.
4
Wait for it
The SDKs poll until the run is
completed or cancelled. With HTTP, poll GET /v1/test_runs/{test_run_id}, or get the test_run.completed webhook.5
Read the results
Response
Start states
A start state is the machine each try begins from. Images (kind: "image", such as win10-22h2) are a runner’s base OS install: read-only, and every org sees them. Your own start states (kind: "state", ids ss_…) are machines made from an image or from another start state, then saved. A case uses one through its start_state, and its setup_script runs on top.
Use a start state for setup that’s slow or can’t be scripted: an app that has to be installed and signed in, a printer that has to be added, a problem only a person can create by clicking. There are three ways to make one with POST /v1/start_states:
start_state.ready or start_state.failed is sent to your webhook when a build ends.
To set one up by hand, or to change a saved one, open its machine:
POST /v1/start_states/{start_state_id}/openboots a machine from its current version (status: "open").GET /v1/start_states/{start_state_id}/screenis its screen as a PNG.POST /v1/start_states/{start_state_id}/inputsends{ "events": [...] }:{ "type": "click", "x", "y", "button"? },double_click,move,{ "type": "scroll", "x", "y", "dy" },{ "type": "type", "text" }and{ "type": "key", "keys": ["CTRL", "ALT", "DELETE"] }.POST /v1/start_states/{start_state_id}/savekeeps the machine as the nextversion(saving, thenready).POST …/closediscards the changes; a hand-made start state that was never saved is deleted.
DELETE answers 409 with code: "in_use" and the cases in test_case_ids.
Runners
Machines live with runners: a small process next to a hypervisor that makes, runs and saves test machines. It connects out to GuidingHand, so it needs no inbound ports. GuidingHand’s cloud runners serve every org with images such aswin10-22h2 (runs_on: "vm"). You can also run your own, for hardware or images of your own: a lab of real laptops, a Windows build with your software installed, a machine inside your network. Your runners serve only your org, and your attempts go to them first.
-
Register one with
POST /v1/runners { "name": "lab-1" }(admin). The response has itstoken(gh_rn_…), shown once. -
Start the runner with that token on a Linux host with KVM (a bare-metal server, or a cloud VM with nested virtualization). It says which images it has, its
labelsand how many machines it runs at once (slots):The runner (Node 22, no dependencies) comes with a QEMU/KVM provider and the scripts that build a Windows 10 image. To run your own, contact us for the package and the setup guide. -
Give cases that need it a matching
runs_on, for example"runs_on": "lenovo"for a runner labelledlenovo.
GET /v1/runners lists your runners and the cloud ones, with status (online or offline), busy and slots. DELETE /v1/runners/{runner_id} revokes the token and disconnects it.
Permissions
Anyone in the org can read test cases, sets, runs, attempts, start states and runners, and start, cancel and rerun runs. Creating and changing test cases, test sets, start states and runners, and deleting runs, needs the admin role or an API key.In the SDKs
Every endpoint is in the API reference, under Test cases, Test sets, Test runs, Start states and Runners.