Skip to content

feat(evals): add agent context eval suite #8427

Description

@sudoKrishna

Problem

The agent decides what the model sees from conversation memory, but nothing
evaluates that end to end. Unit tests cover the windowing functions
(selectConversationTokenWindow, selectConversationContextGroups); nothing
covers whether the assembled provider request actually contains prior history,
in order, with the system prompt preserved, scoped to the right conversation. A
regression in message assembly is invisible.

Proposal

Add an agent-context eval suite that drives the real Agent block through the
DAGExecutor with conversation memory on, stubs the memory read per conversation
id, and asserts the provider request.

  • Location: apps/sim/evals/agent-context/
  • Reuse the executor harness: a case is an executor case plus agent.memory.
  • Checks: all history reached the provider, history precedes the new user prompt,
    the system prompt survived, and the configured conversation id was used.

Acceptance criteria

  • 2+ context scenarios run through the real executor
  • Asserts the assembled provider request (history, order, system prompt,
    conversation id)
  • A wrong conversation id fails the run (verified locally)
  • Results appear in a report and the README documents the suite

Follow-up

Windowing inside the memory service (sliding window size/tokens) and the
retrieval tool are the next layer; this suite covers the assembly the model sees.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    featureNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions