Problem
The agent decides what the model sees from conversation memory, but nothing
evaluates that end to end. Unit tests cover the windowing functions
(selectConversationTokenWindow, selectConversationContextGroups); nothing
covers whether the assembled provider request actually contains prior history,
in order, with the system prompt preserved, scoped to the right conversation. A
regression in message assembly is invisible.
Proposal
Add an agent-context eval suite that drives the real Agent block through the
DAGExecutor with conversation memory on, stubs the memory read per conversation
id, and asserts the provider request.
- Location:
apps/sim/evals/agent-context/
- Reuse the executor harness: a case is an executor case plus
agent.memory.
- Checks: all history reached the provider, history precedes the new user prompt,
the system prompt survived, and the configured conversation id was used.
Acceptance criteria
Follow-up
Windowing inside the memory service (sliding window size/tokens) and the
retrieval tool are the next layer; this suite covers the assembly the model sees.
Problem
The agent decides what the model sees from conversation memory, but nothing
evaluates that end to end. Unit tests cover the windowing functions
(
selectConversationTokenWindow,selectConversationContextGroups); nothingcovers whether the assembled provider request actually contains prior history,
in order, with the system prompt preserved, scoped to the right conversation. A
regression in message assembly is invisible.
Proposal
Add an agent-context eval suite that drives the real Agent block through the
DAGExecutorwith conversation memory on, stubs the memory read per conversationid, and asserts the provider request.
apps/sim/evals/agent-context/agent.memory.the system prompt survived, and the configured conversation id was used.
Acceptance criteria
conversation id)
Follow-up
Windowing inside the memory service (sliding window size/tokens) and the
retrieval tool are the next layer; this suite covers the assembly the model sees.