Skip to content

feat(evals): add adversarial agent tool-use scenarios - #8513

Open
sudoKrishna wants to merge 9 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-adversarial-scenarios
Open

sudoKrishna wants to merge 9 commits into
simstudioai:mainfrom
sudoKrishna:feat/evals-adversarial-scenarios

Conversation

@sudoKrishna

Copy link
Copy Markdown

Summary

Five cases that stress where models actually fail, so the suite measures more
than easy tool calls. Scripted expectations keep them deterministic in CI; the
same cases run in the live and model-comparison suites.

Stacked on #8409 (the eval harness) — base branch is feat/agent-tool-use-evals.

Closes #8512

Scenarios added

Scenario Failure mode
no-tool-needed calls a tool when none is needed
disambiguate-similar-tools picks current weather for a forecast question
empty-result-no-hallucination invents a policy from an empty result
long-chain-dependency loses order across four dependent tools
near-duplicate-names picks get_user over get_user_settings

Test plan

  • bun run test:evals → 17/17 (13 loop + 4 executor), report regenerated
  • New cases appear in the report
  • Scripted expectations deterministic; no provider key
  • bun run check:test-patterns passes
  • Live run to see which cases DeepSeek fails
  • Full bun run type-check — run in CI

Follow-up

Run the live and comparison suites to see real pass rates on these cases; a
failure here is a candidate for a harness or prompt fix, not a flaky test.

Add a deterministic eval layer for the agent harness. Scenarios script the
OpenAI-compatible streaming tool loop with model turns and stub tool results,
then score tool selection, planning, retrieval, and recovery without a
provider key.

- apps/sim/evals/agent-tool-use: 8 scenarios, scoring, JSON+Markdown report
- `bun run test:evals` from apps/sim runs the suite and writes the report
- picked up by the normal vitest run so a regression fails CI
- README documents the contract and how to add a case
Replay the same scenarios against a real model. The model is the only thing
that changes: runScenario now takes an optional completion transport and a
live mode that relaxes exact assertions (ordered subsequence, minimum
successes) and skips scripted-only recovery cases.

- live.ts: OpenAI-compatible transport + DeepSeek factory
- agent-tool-use.live.test.ts: K trials per scenario, gated on
  EVAL_LIVE=1 and DEEPSEEK_API_KEY, never runs in CI
- live report with pass rates, avg iterations, latency, failed checks
- test:evals:live script and README knobs
…ve mode

The first live DeepSeek run exposed brittle assertions, not harness bugs:
the model chained the tools correctly but the checks were case-sensitive and
required an internal order id. Match the retrieved value case-insensitively
and let live runs accept the grounded status rather than the internal id.
Add an executor-level harness: a real Start -> Agent workflow on DAGExecutor,
with only executeProviderRequest mocked at the provider boundary. This covers
agent-block input wiring, variable resolution from Start outputs, and executor
run/error handling, which the direct loop harness cannot see.

- executor-harness.ts: workflow builder + runExecutorScenario
- shares the scorer (scoreExpectations) and report with the loop suite
- two scenarios: Start->Agent output, and <start.message> resolution
- README documents adding an executor-level scenario
Add executor-retries-failed-block: the first provider call rejects, the
Agent block has retry enabled, and the executor replays it. The run must
complete with the second response. Verifies providerCalls === 2, and fails
without the retry policy (checked locally: expected 2, got 1).
Add executor-falls-back-to-secondary-model: the primary call rejects, the
Agent block has a fallback model, and the handler serves the answer from
gpt-4o-mini. Asserts providerCalls === 2 and lastRequestModel, and fails
without the fallback row (checked locally: got gpt-4o, run errored).
Five cases that stress where models tend to fail: answering with no tool,
disambiguating near-identical tools, not inventing an answer from an empty
tool result, running a four-tool dependency chain, and picking settings over
a near-duplicate profile tool. Scripted expectations keep them deterministic;
the same cases run live.
@sudoKrishna
sudoKrishna requested a review from a team as a code owner October 1, 2026 13:33
@vercel

vercel Bot commented Oct 1, 2026

Copy link
Copy Markdown

@sudoKrishna is attempting to deploy a commit to the Sim Team on Vercel.

A member of the Team first needs to authorize it.

@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 4/5

[Medium risk] Adds evaluation suite for the agent tool-use loop.

The evaluation changes have non-blocking measurement issues, but the repository’s absolute-import requirement must be satisfied before merging.

Findings

  1. P2 Live results ignore arguments ▶
  2. P2 Reported calls count as executed ▶
  3. P2 Live command silently skips ▶
  4. P2 Relative imports violate requirement ▶

Summary

Adds scripted and opt-in live agent tool-use evaluations, an executor-level suite, shared scoring, and JSON/Markdown reports.

  • Live stub responses do not validate tool arguments, allowing dependency errors to pass.
  • Executor results count provider-reported calls as successful executions.
  • An explicit live run can silently skip when its API key is absent.
  • New sibling imports need to follow the repository’s absolute-import requirement.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  S[Scenario] --> H[Loop or executor harness]
  H --> M[Scripted or live model / mocked provider]
  M --> T[Tool calls and responses]
  T --> C[Shared expectation scorer]
  C --> R[JSON and Markdown reports]
Loading

Reviews (1) · Last reviewed commit: "feat(evals): add adversarial agent tool-..."

toolsMockFns.mockExecuteTool.mockImplementation(
async (toolId: string, params: Record<string, unknown>): Promise<ToolResponse> => {
const startedAt = Date.now()
const response = resultQueues.get(toolId)?.shift() ?? { success: true, output: {} }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Live results ignore arguments The live stub returns a scripted result based only on the tool name; it records the model’s arguments but does not check them. A model can call read_file with the wrong path or carry the wrong order ID through the four-tool chain and still receive the expected result. That makes a passing trial unreliable as a measure of dependency handling.

Comment on lines +161 to +170
const toolCalls: ScoredToolCall[] = rawToolCalls.map((call) => ({
name: typeof call.name === 'string' ? call.name : 'unknown',
success: true,
}))
const toolInvocations: EvalToolInvocation[] = rawToolCalls.map((call) => ({
name: typeof call.name === 'string' ? call.name : 'unknown',
arguments: (call.arguments ?? {}) as Record<string, unknown>,
success: true,
durationMs: typeof call.duration === 'number' ? call.duration : 0,
}))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Reported calls count as executed This harness marks every tool call in the mocked provider response as successful. The non-streaming Agent path passes those calls into its output without executing them, so executor-agent-runs can pass its successful-tool-call check even though search_docs never ran. The report consequently overstates what this scenario verifies.

* nondeterministic. The report carries pass rates, not a single boolean. Set
* `EVAL_MIN_PASS_RATE` (0–1) to turn a pass-rate floor into a failing gate.
*/
const LIVE = process.env.EVAL_LIVE === '1' && Boolean(process.env.DEEPSEEK_API_KEY)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Live command silently skips If someone runs test:evals:live without DEEPSEEK_API_KEY, this condition skips the entire live suite and prevents the report from being written. The command can then finish successfully despite measuring nothing, making an explicit live run look green when it did not run.

@@ -0,0 +1,592 @@
import type { AgentToolUseScenario, EvalToolDefinition } from './types'

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Relative imports violate requirement This new file imports ./types rather than using the @/evals/agent-tool-use/... alias. The Sim import directive requires absolute imports and prohibits relative imports. The same pattern appears in report.ts, harness.ts, and executor-harness.ts; the repository requirement must be satisfied before merging.

Context Used: Import patterns for the Sim application (source)

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Two live failures were eval design, not model failure:
- empty-result-no-hallucination rejected valid 'didn't find' / 'wasn't able
  to find' phrasing. Broaden the grounding check.
- near-duplicate-names required a userId the prompt never gave, so the model
  reasonably asked for it. Put the id in the prompt and the scripted call.
- long-chain-dependency: the prompt never gave a userId, so the model asked
  or skipped the profile step. Provide u-42 and let live runs require the
  three downstream calls rather than the exact four-step sequence.
- near-duplicate-names: one live trial called both tools; that is over-calling,
  not wrong-tool selection. Drop the forbidden-tool assertion in live mode.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant