> ## Documentation Index
> Fetch the complete documentation index at: https://docs.guidinghand.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evals

> Measure how often an agent fixes a problem: test cases on fresh machines, run N times each against your agents, scored by the machine's end state.

An eval tells you how often an agent fixes a problem before a customer meets it. You describe the problem once as a **test case**. For every try, GuidingHand makes a fresh machine, breaks it the way the case says, and runs a real GuidingHand task on it with your agent: the same agent loop, guardrails, recording and replay as a customer's. A **simulated customer** answers the agent's questions. Then the machine itself is checked.

Group cases into a **test set**, **run** the set against one or more agents with a few tries per case, and compare pass rates side by side. Every try is an **attempt** with its own replay.

Everything here is in the `/v1` API, both SDKs and the console's **Evals** tab.

## Test cases

A test case is a problem on a machine and a way to tell it's fixed:

```json theme={null}
{
  "object": "test_case",
  "test_case_id": "tc_8Kq2mZ0aLp4x",
  "key": "DSP-01",
  "name": "Apps switched to dark mode",
  "prompt": "All my windows suddenly went black. How do I get the white back?",
  "start_state": "win10-22h2",
  "setup_script": "Set-ItemProperty -Path HKCU:\\Software\\Microsoft\\Windows\\CurrentVersion\\Themes\\Personalize -Name AppsUseLightTheme -Value 0",
  "success": {
    "check_script": "(Get-ItemProperty HKCU:\\Software\\Microsoft\\Windows\\CurrentVersion\\Themes\\Personalize).AppsUseLightTheme -eq 1",
    "rubric": null
  },
  "customer": {
    "persona": "Retired teacher. Not technical. Short, polite answers.",
    "facts": { "Wi-Fi network": "HomeNet", "Wi-Fi password": { "secret": "blue-kite-42" } },
    "clicks_admin_prompts": true,
    "does_steps": false
  },
  "operator": { "approvals": "approve" },
  "limits": { "minutes": 10, "steps": null },
  "runs_on": "vm",
  "category": "Display & graphics",
  "tags": ["uac"],
  "metadata": {},
  "created_at": "2026-09-30T15:03:36.037Z",
  "updated_at": "2026-09-30T15:03:36.037Z"
}
```

| Field | What it does |
| - | - |
| `key` | Your own id for the case, unique in the org: 1 to 64 of `A-Za-z0-9._-`, not starting with `tc_`. Every `{test_case_id}` in the API also takes it, so you can use your benchmark's ids. |
| `name` | Up to 200 characters. |
| `prompt` | What the customer says to the agent: the task's prompt. Up to 8,000 characters. |
| `start_state` | The machine each try starts from: an image such as `win10-22h2`, or a [start state](#start-states) of your own. |
| `setup_script` | Runs at the start of every try, elevated, as the signed-in user. Use it to break the machine. Up to 100,000 characters. |
| `success` | How to tell it's fixed: a `check_script`, a `rubric`, or both. See [below](#how-a-try-is-scored). |
| `customer` | The [simulated customer](#the-simulated-customer). `null` uses the test set's, else a built-in default. |
| `operator` | How the agent's approval requests are decided during the test: `approve` or `deny`. |
| `limits` | `minutes` (1 to 60, default 10): the task is stopped after this long. `steps` (1 to 500): model turns, or `null` for the agent's own `guardrails.max_steps`. |
| `runs_on` | A runner label the case needs: `vm` for any virtual machine, or one of your [runners'](#runners) labels, such as `lenovo` for real hardware. |
| `category`, `tags` | For filtering and for the breakdowns in results. Up to 50 tags of up to 64 characters. |

Create cases with `POST /v1/test_cases`, change them with `PATCH /v1/test_cases/{test_case_id}` (only the fields you send change), list them with `GET /v1/test_cases?tag=&category=&test_set_id=&q=`, and delete them with `DELETE /v1/test_cases/{test_case_id}`.

To load a whole benchmark, `POST /v1/test_cases/batch` creates or updates up to 500 cases at once, matched by `key`, and with `test_set_id` adds them to that set. Run it again after editing your file: cases with a known key are updated, new ones are created. It returns `{ "data", "created", "updated" }`.

## How a try is scored

The machine's end state decides. The agent saying "Done." never counts.

* A **check script** runs on the machine (elevated, as the signed-in user) and prints `True` once the problem is fixed. It runs twice: before the task, where it must print `False` (proof the setup really broke the machine), and after it. For a case that asks a question rather than fixes something, the script reads the agent's final answer from the `GH_AGENT_ANSWER` environment variable. `GH_CHECK_PHASE` says which run it is (`before` or `after`): a case where nothing is broken beforehand, like a risky request the agent should refuse, prints `False` when it's `before`.
* A **rubric** says what a fix looks like in plain language. A model judges it from the recorded conversation and the first and last screens, and gives a `reason`.
* A case can have both. An attempt passes only when each one it has passes.

Each attempt ends with an `outcome`:

| `outcome` | Means |
| - | - |
| `passed` | The check printed `True` and the rubric passed (whichever the case has). |
| `failed` | It didn't. A task that failed or hit its time limit is still checked: the machine may be fixed anyway. |
| `error` | Something on GuidingHand's side: the machine didn't start, the app didn't connect, a script crashed, or the check printed something other than `False` before the task. Errors never count against the agent, and are retried automatically up to 2 times. |

## The simulated customer

During a test, nobody sits at the machine. A model plays the customer: it gets the case's `persona`, its `facts`, the agent's question and the machine's screen, and answers the way that person would. It never sees the expected fix or the check.

* `facts` are what the customer knows, by name (up to 50). A value written as `{ "secret": "…" }` is never shown to the model: it only knows the fact's name, and can type the value into the field that has focus, the way a person types a password.
* `clicks_admin_prompts: true` lets it click **Yes** on an admin (UAC) prompt when the agent asks.
* `does_steps: false` makes it hand steps back ("Could you do that for me?") when the agent asks it to do the fix itself.

The simulated customer answers every question the agent asks (as your team, when the agent keeps its questions for your team). Approvals are decided by the case's `operator.approvals`. Simulated answers are marked `simulated` in the task's events, and each attempt lists the customer's turns in `customer`.

## Test sets

A test set is a list of cases to run together, with a default number of tries (`repeat`, 1 to 20) and a default `customer` for its cases that have none.

```json theme={null}
{
  "object": "test_set",
  "test_set_id": "ts_3nV0qR7cWb1e",
  "name": "Windows 10 basics",
  "description": "",
  "repeat": 3,
  "customer": null,
  "cases": [
    { "test_case_id": "tc_8Kq2mZ0aLp4x", "key": "DSP-01", "repeat": null },
    { "test_case_id": "tc_Wb1e3nV0qR7c", "key": "NET-04", "repeat": 5 }
  ],
  "case_count": 2,
  "metadata": {},
  "created_at": "2026-09-30T15:04:10.511Z",
  "updated_at": "2026-09-30T15:04:10.511Z"
}
```

`POST /v1/test_sets/{test_set_id}/cases` adds cases (or changes the `repeat` of ones already in it), `DELETE /v1/test_sets/{test_set_id}/cases/{test_case_id}` takes one out, and `PATCH` with `cases` replaces the whole list. Deleting a set keeps its cases and its runs.

## Runs and attempts

`POST /v1/test_runs` runs a test set (`test_set_id`) or a list of cases (`test_case_ids`) against each of `agents` (default `["default"]`). Tries per case are the first of:

1. `repeat_by_case[case]` in the request,
2. `repeat` in the request,
3. the case's `repeat` in the set,
4. the set's `repeat`,
5. 1.

A run has at most 5,000 attempts. It freezes a copy of each case, and of its start state's version, so editing a case later doesn't change a run that already started.

```json theme={null}
{
  "object": "test_run",
  "test_run_id": "tr_Vb7sT2Lc0aQ3",
  "name": "Nightly",
  "test_set_id": "ts_3nV0qR7cWb1e",
  "agents": ["default", "windows-support"],
  "status": "completed",
  "progress": { "total": 16, "queued": 0, "running": 0, "passed": 11, "failed": 4, "error": 1 },
  "pass_rate": { "default": 0.57, "windows-support": 0.88 },
  "metadata": {},
  "created_at": "2026-09-30T15:05:00.000Z",
  "started_at": "2026-09-30T15:05:02.114Z",
  "finished_at": "2026-09-30T15:41:37.902Z"
}
```

A run is `queued`, then `running`, then `completed` once none of its attempts are left (or `cancelled`). `test_run.completed` is sent to your [webhook](/guides/webhooks) when it ends.

* `GET /v1/test_runs/{test_run_id}/results` has a row per case and agent (`tries`, `passed`, `failed`, `error`, `all_passed`, `median_seconds`, `cost_cents`) and breakdowns `by_agent`, `by_tag` and `by_category`. It counts the last attempt of each try, so an error that was retried counts as its retry.
* `GET /v1/test_runs/{test_run_id}/attempts?test_case_id=&agent_id=&outcome=` lists the attempts. Each has the check's output before and after (`check`), the judge's verdict and reason (`judge`), the simulated customer's turns (`customer`), its `metrics` (active seconds, steps, questions, approvals, cost) and the `replay_url` of its task.
* `GET /v1/attempts/{attempt_id}` is one attempt, with the copy of the test case it ran.
* `POST /v1/test_runs/{test_run_id}/cancel` stops a run. `POST /v1/test_runs/{test_run_id}/rerun` with `{ "only": "failed" }` (or `"error"`, or `"all"`) starts a new run of those (case, agent) pairs.

Eval tasks are billed like any task. They don't count toward the org's limit on tasks running at once: the runners' slots limit them instead.

## Quick start

Create a case, put it in a set, run the set against two agents, wait, and read the results.

<Steps>
  <Step title="Create a test case">
    <CodeGroup>
      ```bash cURL theme={null}
      curl -X POST https://guidinghand.ai/v1/test_cases \
        -H "Authorization: Bearer $GUIDINGHAND_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{
          "key": "DSP-01",
          "name": "Apps switched to dark mode",
          "prompt": "All my windows suddenly went black. How do I get the white back?",
          "start_state": "win10-22h2",
          "setup_script": "Set-ItemProperty -Path HKCU:\\Software\\Microsoft\\Windows\\CurrentVersion\\Themes\\Personalize -Name AppsUseLightTheme -Value 0",
          "success": { "check_script": "(Get-ItemProperty HKCU:\\Software\\Microsoft\\Windows\\CurrentVersion\\Themes\\Personalize).AppsUseLightTheme -eq 1" },
          "customer": { "persona": "Retired teacher. Not technical. Short, polite answers." },
          "tags": ["display"]
        }'
      ```

      ```javascript JavaScript theme={null}
      import GuidingHand from 'guidinghand';

      const gh = new GuidingHand(); // reads GUIDINGHAND_API_KEY

      const key = String.raw`HKCU:\Software\Microsoft\Windows\CurrentVersion\Themes\Personalize`;
      await gh.testCases.create({
        key: 'DSP-01',
        name: 'Apps switched to dark mode',
        prompt: 'All my windows suddenly went black. How do I get the white back?',
        start_state: 'win10-22h2',
        setup_script: `Set-ItemProperty -Path ${key} -Name AppsUseLightTheme -Value 0`,
        success: { check_script: `(Get-ItemProperty ${key}).AppsUseLightTheme -eq 1` },
        customer: { persona: 'Retired teacher. Not technical. Short, polite answers.' },
        tags: ['display'],
      });
      ```

      ```python Python theme={null}
      from guidinghand import GuidingHand

      client = GuidingHand()  # reads GUIDINGHAND_API_KEY

      key = r"HKCU:\Software\Microsoft\Windows\CurrentVersion\Themes\Personalize"
      client.test_cases.create(
          "Apps switched to dark mode",
          "All my windows suddenly went black. How do I get the white back?",
          "win10-22h2",
          {"check_script": f"(Get-ItemProperty {key}).AppsUseLightTheme -eq 1"},
          key="DSP-01",
          setup_script=f"Set-ItemProperty -Path {key} -Name AppsUseLightTheme -Value 0",
          customer={"persona": "Retired teacher. Not technical. Short, polite answers."},
          tags=["display"],
      )
      ```
    </CodeGroup>
  </Step>

  <Step title="Put it in a test set">
    <CodeGroup>
      ```bash cURL theme={null}
      curl -X POST https://guidinghand.ai/v1/test_sets \
        -H "Authorization: Bearer $GUIDINGHAND_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{ "name": "Windows 10 basics", "repeat": 3, "cases": [{ "test_case_id": "DSP-01" }] }'
      ```

      ```javascript JavaScript theme={null}
      const set = await gh.testSets.create({ name: 'Windows 10 basics', repeat: 3, cases: [{ test_case_id: 'DSP-01' }] });
      ```

      ```python Python theme={null}
      test_set = client.test_sets.create("Windows 10 basics", repeat=3, cases=[{"test_case_id": "DSP-01"}])
      ```
    </CodeGroup>
  </Step>

  <Step title="Run it against your agents">
    Three tries per case for each of two agents: six attempts.

    <CodeGroup>
      ```bash cURL theme={null}
      curl -X POST https://guidinghand.ai/v1/test_runs \
        -H "Authorization: Bearer $GUIDINGHAND_API_KEY" \
        -H "Content-Type: application/json" \
        -d '{ "test_set_id": "ts_3nV0qR7cWb1e", "agents": ["default", "windows-support"], "name": "Nightly" }'
      ```

      ```javascript JavaScript theme={null}
      const run = await gh.testSets.run(set.test_set_id, { agents: ['default', 'windows-support'], name: 'Nightly' });
      ```

      ```python Python theme={null}
      run = client.test_sets.run(test_set["test_set_id"], agents=["default", "windows-support"], name="Nightly")
      ```
    </CodeGroup>
  </Step>

  <Step title="Wait for it">
    The SDKs poll until the run is `completed` or `cancelled`. With HTTP, poll `GET /v1/test_runs/{test_run_id}`, or get the `test_run.completed` webhook.

    <CodeGroup>
      ```bash cURL theme={null}
      curl https://guidinghand.ai/v1/test_runs/tr_Vb7sT2Lc0aQ3 \
        -H "Authorization: Bearer $GUIDINGHAND_API_KEY"
      ```

      ```javascript JavaScript theme={null}
      const done = await gh.testRuns.wait(run.test_run_id, { pollInterval: 10_000 }); // { timeout } in ms to give up
      console.log(done.status, done.pass_rate); // completed { default: 0.67, 'windows-support': 1 }
      ```

      ```python Python theme={null}
      done = client.test_runs.wait(run["test_run_id"], poll_interval=10)  # timeout= in seconds to give up
      print(done["status"], done["pass_rate"])
      ```
    </CodeGroup>
  </Step>

  <Step title="Read the results">
    <CodeGroup>
      ```bash cURL theme={null}
      curl https://guidinghand.ai/v1/test_runs/tr_Vb7sT2Lc0aQ3/results \
        -H "Authorization: Bearer $GUIDINGHAND_API_KEY"

      # the failed tries, with their replays
      curl "https://guidinghand.ai/v1/test_runs/tr_Vb7sT2Lc0aQ3/attempts?outcome=failed" \
        -H "Authorization: Bearer $GUIDINGHAND_API_KEY"
      ```

      ```javascript JavaScript theme={null}
      const results = await gh.testRuns.results(run.test_run_id);
      for (const row of results.data) console.log(row.key, row.agent_id, `${row.passed}/${row.tries}`, row.all_passed);

      for await (const a of gh.testRuns.listAllAttempts(run.test_run_id, { outcome: 'failed' })) {
        console.log(a.key, a.agent_id, a.check?.after, a.judge?.reason, a.replay_url);
      }
      ```

      ```python Python theme={null}
      results = client.test_runs.results(run["test_run_id"])
      for row in results["data"]:
          print(row["key"], row["agent_id"], f"{row['passed']}/{row['tries']}", row["all_passed"])

      for a in client.test_runs.list_all_attempts(run["test_run_id"], outcome="failed"):
          print(a["key"], a["agent_id"], (a["check"] or {}).get("after"), (a["judge"] or {}).get("reason"), a["replay_url"])
      ```
    </CodeGroup>

    ```json Response theme={null}
    {
      "data": [
        { "test_case_id": "tc_8Kq2mZ0aLp4x", "key": "DSP-01", "name": "Apps switched to dark mode", "category": null, "tags": ["display"],
          "agent_id": "default", "tries": 3, "passed": 2, "failed": 1, "error": 0, "all_passed": false, "median_seconds": 74, "cost_cents": 300 },
        { "test_case_id": "tc_8Kq2mZ0aLp4x", "key": "DSP-01", "name": "Apps switched to dark mode", "category": null, "tags": ["display"],
          "agent_id": "windows-support", "tries": 3, "passed": 3, "failed": 0, "error": 0, "all_passed": true, "median_seconds": 51, "cost_cents": 200 }
      ],
      "by_agent": { "default": { "tries": 3, "passed": 2, "pass_rate": 0.67 }, "windows-support": { "tries": 3, "passed": 3, "pass_rate": 1 } },
      "by_tag": { "display": { "default": { "tries": 3, "passed": 2 }, "windows-support": { "tries": 3, "passed": 3 } } },
      "by_category": {}
    }
    ```
  </Step>
</Steps>

## Start states

A start state is the machine each try begins from. **Images** (`kind: "image"`, such as `win10-22h2`) are a runner's base OS install: read-only, and every org sees them. Your own start states (`kind: "state"`, ids `ss_…`) are machines made from an image or from another start state, then saved. A case uses one through its `start_state`, and its `setup_script` runs on top.

Use a start state for setup that's slow or can't be scripted: an app that has to be installed and signed in, a printer that has to be added, a problem only a person can create by clicking. There are three ways to make one with `POST /v1/start_states`:

| Send | `method` | What happens |
| - | - | - |
| `script` | `script` | The runner boots `from`, runs the script (elevated, as the signed-in user) and saves the machine. `building`, then `ready` or `failed` with `error`. |
| `instruction` | `instruction` | A GuidingHand task with that prompt (as `agent_id`, default `default`) sets the machine up, then it's saved. The task is in `task_id`. |
| neither | `hand` | The machine boots and stays `open` for someone to set up by hand. |

`start_state.ready` or `start_state.failed` is sent to your webhook when a build ends.

To set one up by hand, or to change a saved one, open its machine:

* `POST /v1/start_states/{start_state_id}/open` boots a machine from its current version (`status: "open"`).
* `GET /v1/start_states/{start_state_id}/screen` is its screen as a PNG.
* `POST /v1/start_states/{start_state_id}/input` sends `{ "events": [...] }`: `{ "type": "click", "x", "y", "button"? }`, `double_click`, `move`, `{ "type": "scroll", "x", "y", "dy" }`, `{ "type": "type", "text" }` and `{ "type": "key", "keys": ["CTRL", "ALT", "DELETE"] }`.
* `POST /v1/start_states/{start_state_id}/save` keeps the machine as the next `version` (`saving`, then `ready`). `POST …/close` discards the changes; a hand-made start state that was never saved is deleted.

Screen and input go through the hypervisor, not the GuidingHand app, so they work on admin prompts and the sign-in screen too. An open machine closes itself after 2 hours without input. The console's **Start states** tab does all of this in the page.

```javascript theme={null}
const state = await gh.startStates.create({ name: 'Printer offline', from: 'win10-22h2', script: 'Stop-Service Spooler; Set-Service Spooler -StartupType Disabled' });
const ready = await gh.startStates.wait(state.start_state_id); // until 'ready' or 'failed'
if (ready.status === 'failed') throw new Error(ready.error ?? 'build failed');
await gh.testCases.update('PRN-02', { start_state: ready.start_state_id });
```

A start state that test cases use can't be deleted: `DELETE` answers `409` with `code: "in_use"` and the cases in `test_case_ids`.

## Runners

Machines live with **runners**: a small process next to a hypervisor that makes, runs and saves test machines. It connects out to GuidingHand, so it needs no inbound ports.

GuidingHand's cloud runners serve every org with images such as `win10-22h2` (`runs_on: "vm"`). You can also run your own, for hardware or images of your own: a lab of real laptops, a Windows build with your software installed, a machine inside your network. Your runners serve only your org, and your attempts go to them first.

1. Register one with `POST /v1/runners { "name": "lab-1" }` (admin). The response has its `token` (`gh_rn_…`), shown once.
2. Start the runner with that token on a Linux host with KVM (a bare-metal server, or a cloud VM with nested virtualization). It says which images it has, its `labels` and how many machines it runs at once (`slots`):

   ```bash theme={null}
   gh-runner --server https://guidinghand.ai --token gh_rn_… --labels lenovo --slots 4
   ```

   The runner (Node 22, no dependencies) comes with a QEMU/KVM provider and the scripts that build a Windows 10 image. To run your own, [contact us](mailto:dev@guidinghand.ai) for the package and the setup guide.
3. Give cases that need it a matching `runs_on`, for example `"runs_on": "lenovo"` for a runner labelled `lenovo`.

`GET /v1/runners` lists your runners and the cloud ones, with `status` (`online` or `offline`), `busy` and `slots`. `DELETE /v1/runners/{runner_id}` revokes the token and disconnects it.

```python theme={null}
runner = client.runners.create("lenovo-lab-1")
print(runner["token"])   # gh_rn_...: store it, it isn't shown again
for r in client.runners.list()["data"]:
    print(r["name"], r["scope"], r["status"], f"{r['busy']}/{r['slots']}", r["labels"])
```

## Permissions

Anyone in the org can read test cases, sets, runs, attempts, start states and runners, and start, cancel and rerun runs. Creating and changing test cases, test sets, start states and runners, and deleting runs, needs the admin role or an API key.

## In the SDKs

| JavaScript | Python | API |
| - | - | - |
| `gh.testCases.list / create / retrieve / update / delete / batch` | `client.test_cases.list / create / retrieve / update / delete / batch` | `/v1/test_cases` |
| `gh.testSets.list / create / retrieve / update / delete / addCases / removeCase / run` | `client.test_sets.… / add_cases / remove_case / run` | `/v1/test_sets` |
| `gh.testRuns.list / create / retrieve / cancel / rerun / delete / results / attempts / wait` | `client.test_runs.… / wait` | `/v1/test_runs` |
| `gh.attempts.retrieve` | `client.attempts.retrieve` | `/v1/attempts/{attempt_id}` |
| `gh.startStates.list / create / retrieve / update / delete / open / save / close / screen / input / wait` | `client.start_states.…` (`create(name, from_, ...)`) | `/v1/start_states` |
| `gh.runners.list / create / delete` | `client.runners.list / create / delete` | `/v1/runners` |

Every endpoint is in the [API reference](/api-reference/introduction), under **Test cases**, **Test sets**, **Test runs**, **Start states** and **Runners**.
