Skip to content

feat(evals): run agent scenarios through the DAGExecutor #8422

Description

@sudoKrishna

Problem

The eval suite drives the provider streaming tool loop directly. It does not
exercise agent-block input wiring, variable resolution, the executor's
retry/fallback policy, or the run/error handling above the loop. A harness change
there can regress without this suite noticing.

Proposal

Add an executor-level harness that builds a serialized workflow with an Agent
block, invokes DAGExecutor.execute, and scores the same checks.

  • Location: apps/sim/evals/agent-tool-use/executor-harness.ts
  • Real Start → Agent workflow; only executeProviderRequest is mocked at the
    provider boundary, so the Agent block handler, variable resolution, and the
    executor run/error handling are real.
  • Reuse @sim/testing factories (createSerializedWorkflow) and the
    @sim/testing mocks already used by
    executor/handlers/agent/agent-handler.test.ts.
  • Reuse the existing scoring/report so executor and loop results land in one
    report.

Acceptance criteria

  • At least one scripted scenario runs through DAGExecutor end to end
  • Asserts the tool calls and the Agent block's final output
  • Results appear in the existing JSON/Markdown report
  • Documented how to add an executor-level scenario

Follow-up

The provider boundary is mocked, so the executor's retry/fallback policy is not
asserted yet. Add a scenario whose first provider call rejects, with a block
retry config, to cover the retry path.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions