Skip to content
143 changes: 143 additions & 0 deletions apps/sim/evals/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,143 @@
# Agent harness evaluations

Measurement for the agent harness — the code that turns a model's tool calls
into executed tools, feeds the results back, and keeps the turn alive when a
tool fails. Unit and integration tests prove the harness handles the cases we
already know about; evals measure whether it still behaves across a suite of
scenarios when the harness changes.

## What runs

The first suite lives in [`agent-tool-use/`](./agent-tool-use) and drives the
real OpenAI-compatible streaming tool loop
(`apps/sim/providers/openai-compat/streaming-tool-loop.ts`) — the loop that
serves OpenAI, DeepSeek, Groq, Cerebras, and the other OpenAI-compatible
providers. The model is **scripted**: each scenario supplies the assistant turns
(tool calls or a final answer) and the result of each tool call. That keeps the
suite deterministic and runnable in CI with no provider key, while the thing
being measured — tool dispatch, result feedback, error recovery — is real
production code.

The suites cover four behaviors:

| Category | What it measures |
| --- | --- |
| `tool-selection` | The loop dispatches the tool the model asked for, including from a set of distractors. |
| `planning` | Multi-turn, dependent and parallel tool calls execute in the right order and all results reach the next turn. |
| `retrieval` | Values returned by a tool survive into the final answer instead of being dropped or invented. |
| `recovery` | Tool errors, unknown tool names, and malformed argument JSON are fed back to the model rather than thrown out of the loop. |

## Run it

From `apps/sim`:

```sh
bun run test:evals
```

The command writes a JSON report and a Markdown summary to
`test-results/evals/agent-tool-use.{json,md}` (gitignored) and fails the process
if any scenario fails. To point the report somewhere else, run Vitest directly:

```sh
EVAL_REPORT_PATH=/tmp/agent-tool-use.json bunx vitest run evals/agent-tool-use
```

The suite is also picked up by the normal `bun run test` run, so a regression
fails CI even without the dedicated command.

## Run against a real model (live)

The same scenarios can be replayed against a live model. This is opt-in and
never runs in CI. DeepSeek is wired first; any OpenAI-compatible provider works
through `createOpenAICompatLiveCompletion` in `live.ts`.

```sh
cd apps/sim
DEEPSEEK_API_KEY=... bun run test:evals:live
```

Useful knobs:

| Variable | Default | Meaning |
| --- | --- | --- |
| `EVAL_TRIALS` | `3` | Runs per scenario. Models are nondeterministic, so results are pass rates. |
| `EVAL_MIN_PASS_RATE` | `0` | When > 0, fail a scenario below this pass rate (0–1). |
| `EVAL_MODEL` | `deepseek-chat` | Model id sent to the provider. |
| `EVAL_TIMEOUT_MS` | `180000` | Per-request timeout. |
| `EVAL_REPORT_PATH` | `test-results/evals/agent-tool-use-live.json` | Report location. |

Live runs relax exact assertions: `toolCallSequence` becomes an ordered
subsequence, `successfulToolCalls` becomes a minimum, and scripted-only cases
(malformed JSON, unknown tool) are skipped. A scenario-level `liveExpect`
overrides the scripted expectation where a real model cannot reproduce it (for
example, an exact retry count). The report is at
`test-results/evals/agent-tool-use-live.{json,md}` with pass rates, average
iterations, latency, and the failed check names.

## Add a case

1. Open [`agent-tool-use/scenarios.ts`](./agent-tool-use/scenarios.ts) and add
an entry to `AGENT_TOOL_USE_SCENARIOS`.
2. Declare the `tools` the model may call and the `script` it produces. A
`tools` turn lists the calls the model emits; an `answer` turn ends the run.
Attach each call's stub `result` (or leave it to default to a successful
empty output).
3. Add the assertions you care about under `expect`: the ordered
`toolCallSequence`, `requiredTools`/`forbiddenTools`, `finalContent`,
`maxIterations`, and tool call counts. Every assertion becomes a named check
in the report.
4. Run `bun run test:evals`.

A scenario is data, not code — there is no harness change needed for a new case.

### Simulating a failure

- **Tool error:** give the call `result: { success: false, error: '...' }`.
- **Unknown tool:** call a `name` that is not in `tools`; the loop returns a
tool-not-found error to the model.
- **Malformed arguments:** set `argumentsJson` to an invalid or non-object JSON
string. The loop must not execute the call and must return the parse error to
the model.

## Executor-level scenarios

[`agent-tool-use/executor-harness.ts`](./agent-tool-use/executor-harness.ts)
runs a case through a real `DAGExecutor`: a Start block → Agent block workflow,
with only the provider boundary (`executeProviderRequest`) mocked. This covers
what the loop harness cannot — agent-block input wiring, variable resolution
from Start outputs, and the executor's run/error handling. Tool dispatch stays
covered by the loop suite.

Add a case to `EXECUTOR_SCENARIOS` in `executor-harness.ts`:

- `workflowInput` is exposed on the Start block; reference an output with
`<start.field>` from the Agent prompt.
- `agent` is the Agent block config (`model`, `systemPrompt`, `userPrompt`).
- `providerResponse` is what the mocked provider returns (`content`,
`toolCalls`, `tokens`).
- `expect` uses the loop's checks plus `resolvedInput` (a substring that must
reach the provider messages), `succeeds` (expected `ExecutionResult.success`),
and `providerCalls` (exact provider call count).
- Set `agent.retry` to exercise the executor's per-block retry policy, or
`agent.fallbackModels` to exercise model fallback. Make the first
`providerResponse` a `reject` and the next one succeeds; assert
`providerCalls` and `lastRequestModel` to prove which path recovered.

Both suites write one report, so executor rows appear alongside loop rows.

## Report shape

`report.json` is machine-readable for dashboards and trend tracking; `report.md`
is the same data as a table. Each result carries the scenario id, pass/fail,
every named check with a failure detail, the final content, the executed tool
invocations, and metrics: iterations, tool call counts (success/error), latency,
model/tool time, first-response time, and token usage.

## Scope and next steps

Two harnesses share one result shape and report: the tool loop and the
`DAGExecutor`. The executor suite covers both recovery paths — block retry
(`executor-retries-failed-block`) and model fallback
(`executor-falls-back-to-secondary-model`). Further expansion (context/memory,
model routing, subagent orchestration) is tracked as follow-up work.
81 changes: 81 additions & 0 deletions apps/sim/evals/agent-tool-use/agent-tool-use.eval.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,81 @@
import {
permissionCheckMock,
permissionCheckMockFns,
} from '@sim/testing/mocks/permission-check.mock'
import { providersMock } from '@sim/testing/mocks/providers.mock'
import { providersConversationHistoryMock } from '@sim/testing/mocks/providers-conversation-history.mock'
import { providersUtilsMock, providersUtilsMockFns } from '@sim/testing/mocks/providers-utils.mock'
import { toolsMock } from '@sim/testing/mocks/tools.mock'
import { workspaceFileSecretProvenanceMock } from '@sim/testing/mocks/workspace-file-secret-provenance.mock'
import { afterAll, beforeEach, describe, expect, it, vi } from 'vitest'
import { EXECUTOR_SCENARIOS, runExecutorScenario } from '@/evals/agent-tool-use/executor-harness'
import { runScenario } from '@/evals/agent-tool-use/harness'
import { writeEvalReport } from '@/evals/agent-tool-use/report'
import { AGENT_TOOL_USE_SCENARIOS } from '@/evals/agent-tool-use/scenarios'
import type { AgentToolUseResult } from '@/evals/agent-tool-use/types'

vi.mock('@/providers/conversation-history', () => providersConversationHistoryMock)
vi.mock('@/tools', () => toolsMock)
vi.mock('@/providers/utils', () => providersUtilsMock)
vi.mock('@/providers', () => providersMock)
vi.mock('@/ee/access-control/utils/permission-check', () => permissionCheckMock)
vi.mock(
'@/lib/uploads/contexts/workspace/workspace-file-secret-provenance',
() => workspaceFileSecretProvenanceMock
)
vi.mock('@/lib/memory/agent-turn-session', () => ({
openAgentTurnSession: vi.fn(async () => undefined),
}))
vi.mock('@/lib/internal/mcp/discover-tools', () => ({
discoverMcpServerToolsAsExecutor: vi.fn(async () => []),
}))
vi.mock('@/lib/internal/custom-tools/read-available-by-id-or-title', () => ({
readAvailableCustomToolByIdOrTitleAsExecutor: vi.fn(async () => undefined),
}))
vi.mock('@/executor/utils/http', () => ({
buildAuthHeaders: vi.fn(async () => ({ 'Content-Type': 'application/json' })),
buildAPIUrl: vi.fn((path: string) => path),
extractAPIErrorMessage: vi.fn(async () => 'request failed'),
}))
vi.mock('@/lib/execution/cancellation', () => ({
subscribeToExecutionCancellation: vi.fn(async () => () => {}),
isExecutionCancelled: vi.fn(async () => false),
}))

const results: AgentToolUseResult[] = []

beforeEach(() => {
permissionCheckMockFns.mockValidateModelProvider.mockResolvedValue(undefined)
providersUtilsMockFns.mockGetProviderFromModel.mockReturnValue('mock-provider')
})

afterAll(() => {
const reportPath = process.env.EVAL_REPORT_PATH
if (reportPath) writeEvalReport(results, reportPath)
})

describe('agent tool-use eval suite', () => {
it.each(AGENT_TOOL_USE_SCENARIOS)('$id: $name', async (scenario) => {
const result = await runScenario(scenario)
results.push(result)

const failed = result.checks.filter((entry) => !entry.passed)
expect(
failed,
failed.map((entry) => `${entry.name}: ${entry.detail}`).join('; ') || undefined
).toEqual([])
})
})

describe('agent executor eval suite', () => {
it.each(EXECUTOR_SCENARIOS)('$id: $name', async (scenario) => {
const result = await runExecutorScenario(scenario)
results.push(result)

const failed = result.checks.filter((entry) => !entry.passed)
expect(
failed,
failed.map((entry) => `${entry.name}: ${entry.detail}`).join('; ') || undefined
).toEqual([])
})
})
83 changes: 83 additions & 0 deletions apps/sim/evals/agent-tool-use/agent-tool-use.live.test.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
import { providersMock } from '@sim/testing/mocks/providers.mock'
import { providersConversationHistoryMock } from '@sim/testing/mocks/providers-conversation-history.mock'
import { providersUtilsMock } from '@sim/testing/mocks/providers-utils.mock'
import { toolsMock } from '@sim/testing/mocks/tools.mock'
import { afterAll, describe, expect, it, vi } from 'vitest'
import { runScenario } from '@/evals/agent-tool-use/harness'
import { createDeepSeekLiveCompletion } from '@/evals/agent-tool-use/live'
import { writeLiveEvalReport } from '@/evals/agent-tool-use/report'
import { AGENT_TOOL_USE_SCENARIOS } from '@/evals/agent-tool-use/scenarios'
import type { AgentToolUseResult, LiveScenarioSummary } from '@/evals/agent-tool-use/types'

vi.mock('@/providers/conversation-history', () => providersConversationHistoryMock)
vi.mock('@/tools', () => toolsMock)
vi.mock('@/providers/utils', () => providersUtilsMock)
vi.mock('@/providers', () => providersMock)

/**
* Live agent tool-use evals. Opt-in only:
*
* EVAL_LIVE=1 DEEPSEEK_API_KEY=... \
* bun run --cwd apps/sim test --mode live evals/agent-tool-use/agent-tool-use.live.test.ts
*
* Each scenario runs `EVAL_TRIALS` times (default 3) because a real model is
* nondeterministic. The report carries pass rates, not a single boolean. Set
* `EVAL_MIN_PASS_RATE` (0–1) to turn a pass-rate floor into a failing gate.
*/
const LIVE = process.env.EVAL_LIVE === '1' && Boolean(process.env.DEEPSEEK_API_KEY)
const TRIALS = Number(process.env.EVAL_TRIALS ?? '3')
const MIN_PASS_RATE = Number(process.env.EVAL_MIN_PASS_RATE ?? '0')
const MODEL = process.env.EVAL_MODEL ?? 'deepseek-chat'
const TIMEOUT_MS = Number(process.env.EVAL_TIMEOUT_MS ?? '180000')

const liveScenarios = AGENT_TOOL_USE_SCENARIOS.filter((scenario) => !scenario.scriptedOnly)
const summaries: LiveScenarioSummary[] = []

afterAll(() => {
if (!LIVE) return
writeLiveEvalReport(
summaries,
process.env.EVAL_REPORT_PATH ?? 'test-results/evals/agent-tool-use-live.json'
)
})

describe.skipIf(!LIVE)('agent tool-use eval suite (live DeepSeek)', () => {
it.each(liveScenarios)(
'$id: $name',
async (scenario) => {
const completion = createDeepSeekLiveCompletion(MODEL)
const results: AgentToolUseResult[] = []

for (let trial = 0; trial < TRIALS; trial++) {
results.push(
await runScenario(scenario, {
completion,
mode: 'live',
model: MODEL,
providerName: 'DeepSeek',
})
)
}

const passed = results.filter((result) => result.passed).length
const passRate = results.length === 0 ? 0 : passed / results.length
summaries.push({
id: scenario.id,
name: scenario.name,
category: scenario.category,
trials: results.length,
passed,
passRate,
results,
})

if (MIN_PASS_RATE > 0) {
expect(
passRate,
`${scenario.id} passed ${passed}/${results.length} trials`
).toBeGreaterThanOrEqual(MIN_PASS_RATE)
}
},
TIMEOUT_MS
)
})
Loading