Skip to content

Resuming a named ChatAgent appends a second system prompt instead of restoring the session #331

Description

@BrennanTM

What happens

With agent_name persistence, a new ChatAgent instance resumes the stored thread by appending rather than restoring. ChatAgent.format_query (chat_agent.py:49-56) seeds a fresh SystemMessage whenever it is handed no prior state, and every launch hands it none: the REPL's AgentHITL.state starts as None (runtime.py:41, :154) and nothing hydrates it from the checkpoint, and a library caller doing the documented ChatAgent(llm=llm, agent_name="proj") has no prior state to pass either. The add_messages reducer then appends the whole block (the new SystemMessage carries a fresh id, so it cannot dedupe), and message counts grow 3, 6, 9 across three "fresh" sessions with one extra identical system prompt each time.

The first request of the resumed session therefore carries the roles system, human, ai, system, human. langchain-anthropic rejects that shape with Received multiple non-consecutive system messages (the same class as #294, which #308 fixed for the deep-review agent), so on Anthropic-backed configs the first chat turn after restarting a named session fails. On OpenAI-style configs the client passes the duplicated prompts through silently. _summarize_context keeps only index 0 as the system message (base.py:833), so the extra prompts are treated as conversation, and when one falls inside the summarized span the summarizer's own request has the same rejected shape. ExecutionAgent is not affected in this way: it injects its prompt per request (execution_agent.py:369-374), so its resumed request stays provider-valid. Neither #308 nor the #297 request harness, which drives single-instance histories only, covers a resume across instances.

Reproduction (no API keys needed):

import os, tempfile
os.environ["URSA_CACHE_HOME"] = tempfile.mkdtemp()  # keep dens out of ~/.cache

from pathlib import Path
from typing import ClassVar

from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
from langchain_core.messages import AIMessage
from ursa.agents.chat_agent import ChatAgent

class RecordingFakeLLM(GenericFakeChatModel):
    calls: ClassVar[list] = []
    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        type(self).calls.append([m.type for m in messages])
        return super()._generate(messages, stop=stop, run_manager=run_manager, **kwargs)
    def bind_tools(self, tools, **kwargs):
        return self

def llm():
    return RecordingFakeLLM(messages=iter([AIMessage("reply")] * 10))

ws = Path(tempfile.mkdtemp()) / "ws"
a = ChatAgent(llm=llm(), workspace=ws, agent_name="proj")
a.invoke(a.format_query("first session")); a.close()
b = ChatAgent(llm=llm(), workspace=ws, agent_name="proj")
r = b.invoke(b.format_query("second session")); b.close()
print(RecordingFakeLLM.calls[-1])  # ['system', 'human', 'ai', 'system', 'human']
assert sum(m.type == "system" for m in r["messages"]) >= 2  # merged, not restored

Where it comes from

The thread id is the same on every launch by design: the REPL pins self.config.thread_id or "ursa" (runtime.py:185) and passes it to the agent (:355), BaseAgent defaults to the same constant (base.py:316), and #183 removed the per-agent suffix precisely so the CLI can pick up checkpoints. Resuming the thread is the intended feature. The defect is that the resume re-seeds the system prompt instead of restoring the conversation. (Line references are as of current main, 9e32d8c.)

Possible fixes, smallest first

  1. Keep the outbound request provider-valid: before the model call, keep the leading system prompt and drop any later SystemMessage from the request, the way fix: keep deep-review role prompts out of the message channel #308 did for deep review. Smallest change, and it also covers the summarizer request.
  2. Make a resumed session restore: have ChatAgent.format_query skip seeding a prompt when the state it receives already holds one, and have the runtime hydrate AgentHITL.state from the checkpoint on launch, so the add_messages reducer dedupes by id. Cleaner, and it touches the CLI runtime.
  3. Both, with 1 as the guard and 2 as the behavior.

I can PR option 1 with a regression case in the #297 request harness that drives a resume across instances, or whichever shape you prefer.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions