Skip to main content
An eval tells you how often an agent fixes a problem before a customer meets it. You describe the problem once as a test case. For every try, GuidingHand makes a fresh machine, breaks it the way the case says, and runs a real GuidingHand task on it with your agent: the same agent loop, guardrails, recording and replay as a customer’s. A simulated customer answers the agent’s questions. Then the machine itself is checked. Group cases into a test set, run the set against one or more agents with a few tries per case, and compare pass rates side by side. Every try is an attempt with its own replay. Everything here is in the /v1 API, both SDKs and the console’s Evals tab.

Test cases

A test case is a problem on a machine and a way to tell it’s fixed:
Create cases with POST /v1/test_cases, change them with PATCH /v1/test_cases/{test_case_id} (only the fields you send change), list them with GET /v1/test_cases?tag=&category=&test_set_id=&q=, and delete them with DELETE /v1/test_cases/{test_case_id}. To load a whole benchmark, POST /v1/test_cases/batch creates or updates up to 500 cases at once, matched by key, and with test_set_id adds them to that set. Run it again after editing your file: cases with a known key are updated, new ones are created. It returns { "data", "created", "updated" }.

How a try is scored

The machine’s end state decides. The agent saying “Done.” never counts.
  • A check script runs on the machine (elevated, as the signed-in user) and prints True once the problem is fixed. It runs twice: before the task, where it must print False (proof the setup really broke the machine), and after it. For a case that asks a question rather than fixes something, the script reads the agent’s final answer from the GH_AGENT_ANSWER environment variable. GH_CHECK_PHASE says which run it is (before or after): a case where nothing is broken beforehand, like a risky request the agent should refuse, prints False when it’s before.
  • A rubric says what a fix looks like in plain language. A model judges it from the recorded conversation and the first and last screens, and gives a reason.
  • A case can have both. An attempt passes only when each one it has passes.
Each attempt ends with an outcome:

The simulated customer

During a test, nobody sits at the machine. A model plays the customer: it gets the case’s persona, its facts, the agent’s question and the machine’s screen, and answers the way that person would. It never sees the expected fix or the check.
  • facts are what the customer knows, by name (up to 50). A value written as { "secret": "…" } is never shown to the model: it only knows the fact’s name, and can type the value into the field that has focus, the way a person types a password.
  • clicks_admin_prompts: true lets it click Yes on an admin (UAC) prompt when the agent asks.
  • does_steps: false makes it hand steps back (“Could you do that for me?”) when the agent asks it to do the fix itself.
The simulated customer answers every question the agent asks (as your team, when the agent keeps its questions for your team). Approvals are decided by the case’s operator.approvals. Simulated answers are marked simulated in the task’s events, and each attempt lists the customer’s turns in customer.

Test sets

A test set is a list of cases to run together, with a default number of tries (repeat, 1 to 20) and a default customer for its cases that have none.
POST /v1/test_sets/{test_set_id}/cases adds cases (or changes the repeat of ones already in it), DELETE /v1/test_sets/{test_set_id}/cases/{test_case_id} takes one out, and PATCH with cases replaces the whole list. Deleting a set keeps its cases and its runs.

Runs and attempts

POST /v1/test_runs runs a test set (test_set_id) or a list of cases (test_case_ids) against each of agents (default ["default"]). Tries per case are the first of:
  1. repeat_by_case[case] in the request,
  2. repeat in the request,
  3. the case’s repeat in the set,
  4. the set’s repeat,
A run has at most 5,000 attempts. It freezes a copy of each case, and of its start state’s version, so editing a case later doesn’t change a run that already started.
A run is queued, then running, then completed once none of its attempts are left (or cancelled). test_run.completed is sent to your webhook when it ends.
  • GET /v1/test_runs/{test_run_id}/results has a row per case and agent (tries, passed, failed, error, all_passed, median_seconds, cost_cents) and breakdowns by_agent, by_tag and by_category. It counts the last attempt of each try, so an error that was retried counts as its retry.
  • GET /v1/test_runs/{test_run_id}/attempts?test_case_id=&agent_id=&outcome= lists the attempts. Each has the check’s output before and after (check), the judge’s verdict and reason (judge), the simulated customer’s turns (customer), its metrics (active seconds, steps, questions, approvals, cost) and the replay_url of its task.
  • GET /v1/attempts/{attempt_id} is one attempt, with the copy of the test case it ran.
  • POST /v1/test_runs/{test_run_id}/cancel stops a run. POST /v1/test_runs/{test_run_id}/rerun with { "only": "failed" } (or "error", or "all") starts a new run of those (case, agent) pairs.
Eval tasks are billed like any task. They don’t count toward the org’s limit on tasks running at once: the runners’ slots limit them instead.

Quick start

Create a case, put it in a set, run the set against two agents, wait, and read the results.
1

Create a test case

2

Put it in a test set

3

Run it against your agents

Three tries per case for each of two agents: six attempts.
4

Wait for it

The SDKs poll until the run is completed or cancelled. With HTTP, poll GET /v1/test_runs/{test_run_id}, or get the test_run.completed webhook.
5

Read the results

Response

Start states

A start state is the machine each try begins from. Images (kind: "image", such as win10-22h2) are a runner’s base OS install: read-only, and every org sees them. Your own start states (kind: "state", ids ss_…) are machines made from an image or from another start state, then saved. A case uses one through its start_state, and its setup_script runs on top. Use a start state for setup that’s slow or can’t be scripted: an app that has to be installed and signed in, a printer that has to be added, a problem only a person can create by clicking. There are three ways to make one with POST /v1/start_states: start_state.ready or start_state.failed is sent to your webhook when a build ends. To set one up by hand, or to change a saved one, open its machine:
  • POST /v1/start_states/{start_state_id}/open boots a machine from its current version (status: "open").
  • GET /v1/start_states/{start_state_id}/screen is its screen as a PNG.
  • POST /v1/start_states/{start_state_id}/input sends { "events": [...] }: { "type": "click", "x", "y", "button"? }, double_click, move, { "type": "scroll", "x", "y", "dy" }, { "type": "type", "text" } and { "type": "key", "keys": ["CTRL", "ALT", "DELETE"] }.
  • POST /v1/start_states/{start_state_id}/save keeps the machine as the next version (saving, then ready). POST …/close discards the changes; a hand-made start state that was never saved is deleted.
Screen and input go through the hypervisor, not the GuidingHand app, so they work on admin prompts and the sign-in screen too. An open machine closes itself after 2 hours without input. The console’s Start states tab does all of this in the page.
A start state that test cases use can’t be deleted: DELETE answers 409 with code: "in_use" and the cases in test_case_ids.

Runners

Machines live with runners: a small process next to a hypervisor that makes, runs and saves test machines. It connects out to GuidingHand, so it needs no inbound ports. GuidingHand’s cloud runners serve every org with images such as win10-22h2 (runs_on: "vm"). You can also run your own, for hardware or images of your own: a lab of real laptops, a Windows build with your software installed, a machine inside your network. Your runners serve only your org, and your attempts go to them first.
  1. Register one with POST /v1/runners { "name": "lab-1" } (admin). The response has its token (gh_rn_…), shown once.
  2. Start the runner with that token on a Linux host with KVM (a bare-metal server, or a cloud VM with nested virtualization). It says which images it has, its labels and how many machines it runs at once (slots):
    The runner (Node 22, no dependencies) comes with a QEMU/KVM provider and the scripts that build a Windows 10 image. To run your own, contact us for the package and the setup guide.
  3. Give cases that need it a matching runs_on, for example "runs_on": "lenovo" for a runner labelled lenovo.
GET /v1/runners lists your runners and the cloud ones, with status (online or offline), busy and slots. DELETE /v1/runners/{runner_id} revokes the token and disconnects it.

Permissions

Anyone in the org can read test cases, sets, runs, attempts, start states and runners, and start, cancel and rerun runs. Creating and changing test cases, test sets, start states and runners, and deleting runs, needs the admin role or an API key.

In the SDKs

Every endpoint is in the API reference, under Test cases, Test sets, Test runs, Start states and Runners.