> ## Documentation Index
> Fetch the complete documentation index at: https://docs.guidinghand.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Compare test runs

> 2 to 10 runs side by side, the first as the baseline: each run’s totals and agents, each case’s results per run and agent, and which (case, agent) pairs improved or regressed. Cases line up by test case and agents by agent id, so runs of a set, reruns and runs with other agents all compare.



## OpenAPI

````yaml /api-reference/openapi.json get /v1/test_runs/compare
openapi: 3.1.0
info:
  title: GuidingHand API
  version: 1.0.0
  description: >-
    Create sessions (an invite link with a code for the person at the computer),
    run tasks on their computer with one of your agents, follow them, answer
    their questions and approvals, and fetch history and recordings. Test your
    agents with evals: test cases on fresh machines, run in test sets, scored by
    a check script or a rubric. Agents are versioned (save a draft, publish it),
    have tools (their own file system, and your HTTP endpoints), and every
    change is in the org’s audit log.
servers:
  - url: https://guidinghand.ai
    description: Production
  - url: https://dev.guidinghand.ai
    description: Development (Stripe test mode)
security:
  - bearerAuth: []
tags:
  - name: Agents
  - name: Agent files
  - name: Agent tools
  - name: Sessions
  - name: Tasks
  - name: Webhooks
  - name: Audit log
  - name: Test cases
  - name: Test sets
  - name: Test runs
  - name: Start states
  - name: Runners
paths:
  /v1/test_runs/compare:
    get:
      tags:
        - Test runs
      summary: Compare test runs
      description: >-
        2 to 10 runs side by side, the first as the baseline: each run’s totals
        and agents, each case’s results per run and agent, and which (case,
        agent) pairs improved or regressed. Cases line up by test case and
        agents by agent id, so runs of a set, reruns and runs with other agents
        all compare.
      operationId: compareTestRuns
      parameters:
        - name: ids
          in: query
          required: true
          description: 2 to 10 test run ids, comma-separated. The first is the baseline.
          schema:
            type: string
          example: tr_a1b2c3d4e5f6,tr_g7h8i9j0k1l2
      responses:
        '200':
          description: OK
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/TestRunComparison'
        '400':
          $ref: '#/components/responses/E400'
        '401':
          $ref: '#/components/responses/E401'
        '404':
          $ref: '#/components/responses/E404'
components:
  schemas:
    TestRunComparison:
      type: object
      description: >-
        Runs side by side. The first run is the baseline; cases line up by test
        case and agents by agent id.
      properties:
        object:
          type: string
          const: test_run_comparison
        baseline:
          type: string
          description: The first run’s id.
        runs:
          type: array
          items:
            type: object
            properties:
              test_run_id:
                type: string
              name:
                type: string
              test_set_id:
                type:
                  - string
                  - 'null'
              status:
                $ref: '#/components/schemas/TestRunStatus'
              agents:
                type: array
                items:
                  type: string
              created_at:
                type: string
                format: date-time
              totals:
                $ref: '#/components/schemas/TestRunTotals'
              by_agent:
                type: object
                additionalProperties:
                  $ref: '#/components/schemas/AgentResults'
        changes:
          type: object
          description: >-
            Per run but the baseline: its (case, agent) pairs scored in both,
            and how their pass rate moved.
          additionalProperties:
            type: object
            properties:
              compared:
                type: integer
              improved:
                type: integer
              regressed:
                type: integer
              unchanged:
                type: integer
        cases:
          type: array
          items:
            type: object
            properties:
              test_case_id:
                type: string
              key:
                type:
                  - string
                  - 'null'
              name:
                type: string
              category:
                type:
                  - string
                  - 'null'
              tags:
                type: array
                items:
                  type: string
              results:
                type: object
                description: >-
                  Per run, per agent. A run without the case (or the agent) has
                  no entry.
                additionalProperties:
                  type: object
                  additionalProperties:
                    type: object
                    properties:
                      tries:
                        type: integer
                      passed:
                        type: integer
                      failed:
                        type: integer
                      error:
                        type: integer
                      median_seconds:
                        type:
                          - number
                          - 'null'
                      price_cents:
                        type: number
              changes:
                type: object
                description: >-
                  Per run, per agent: only the pairs whose pass rate moved from
                  the baseline.
                additionalProperties:
                  type: object
                  additionalProperties:
                    type: string
                    enum:
                      - improved
                      - regressed
    TestRunStatus:
      type: string
      enum:
        - queued
        - running
        - completed
        - cancelled
    TestRunTotals:
      type: object
      properties:
        tries:
          type: integer
          description: The last attempt of each try.
        attempts:
          type: integer
          description: Every attempt, retries included.
        passed:
          type: integer
        failed:
          type: integer
        error:
          type: integer
        pass_rate:
          type:
            - number
            - 'null'
        active_seconds:
          type: integer
        price_cents:
          type: number
        model_cost_cents:
          type: number
        started_at:
          type:
            - string
            - 'null'
          format: date-time
        finished_at:
          type:
            - string
            - 'null'
          format: date-time
        wall_seconds:
          type:
            - integer
            - 'null'
          description: From the run’s start to its end (or now).
    AgentResults:
      type: object
      description: >-
        An agent’s tries in a run. `pass_rate`: passed / (passed + failed);
        errors are ours and never count against it.
      properties:
        tries:
          type: integer
        passed:
          type: integer
        failed:
          type: integer
        error:
          type: integer
        pass_rate:
          type:
            - number
            - 'null'
        active_seconds:
          type: integer
          description: Active time of its tries, retried errors included.
        median_seconds:
          type:
            - number
            - 'null'
          description: Median active time of its scored tries.
        p90_seconds:
          type:
            - number
            - 'null'
          description: 90th percentile active time of its scored tries.
        avg_steps:
          type:
            - number
            - 'null'
          description: Average steps per scored try.
        price_cents:
          type: number
          description: >-
            What its tries cost at the agent’s list price (what a customer would
            pay), retried errors included.
        model_cost_cents:
          type: number
          description: What the model provider charged for its tries, from token counts.
    Error:
      type: object
      required:
        - error
      properties:
        error:
          type: object
          required:
            - type
            - message
          properties:
            type:
              type: string
              enum:
                - invalid_request
                - authentication
                - payment_required
                - permission
                - not_found
                - conflict
                - rate_limit
                - server_error
            message:
              type: string
            code:
              type: string
              description: >-
                Why it was refused, when there’s more to say than the status.
                Starting a task: `insufficient_balance`,
                `auto_topup_limit_reached` or `card_needs_attention` (402: the
                balance can’t cover a minute; add funds, raise the auto top-up
                limit, or update the card), `concurrency_limit` (429),
                `model_unavailable` (503: the agent’s `model` can’t run on this
                server; choose another). On `/respond`: `already_answered`
                (someone answered or decided first: see `answered_by`) or
                `not_pending` (that question or approval isn’t open any more).
                The GuidingHand app hears two more, which the API never returns:
                `customer_answers_off` and `customer_approvals_off` (the agent
                keeps its questions or approvals for your team). Deleting a
                start state that test cases use: `in_use` (409, with
                `test_case_ids`).
            answered_by:
              type: string
              enum:
                - customer
                - operator
                - timeout
              description: >-
                With `already_answered`: who answered or decided first
                (`customer`: the person at the computer; `timeout`: nobody
                decided an approval within the agent’s
                `approval_timeout_minutes`, so it was denied).
            test_case_ids:
              type: array
              items:
                type: string
              description: 'With `in_use`: the test cases that use the start state.'
  responses:
    E400:
      description: The request is invalid.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
    E401:
      description: Missing or invalid API key.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
    E404:
      description: Not found in this org.
      content:
        application/json:
          schema:
            $ref: '#/components/schemas/Error'
  securitySchemes:
    bearerAuth:
      type: http
      scheme: bearer
      description: An org API key (`gh_live_…`) from Settings → API keys in the console.

````