Skip to content

fix: wait for an in-progress reply before sending after a reload - #59

Open
asasemahmed wants to merge 1 commit into
CopilotKit:mainfrom
asasemahmed:fix/wait-for-active-run
Open

asasemahmed wants to merge 1 commit into
CopilotKit:mainfrom
asasemahmed:fix/wait-for-active-run

Conversation

@asasemahmed

Copy link
Copy Markdown
Contributor

What changed

Reloading the app while a reply is still running, then sending a message, fails with:

Thread <uuid> is locked

A Retry response button appears next to it, and retrying would ask for the earlier reply again.

The cause: hydrate() marks the conversation loaded as soon as replayed messages arrive, while connectAgent is still attached to the running reply. The next turn is therefore sent while the server still holds the thread lock. CopilotKit reports the refusal as agent_thread_locked, then again as agent_run_failed.

Now:

  • New turns stay queued until connectAgent returns. In connect mode it returns once the thread is idle. So after a reload, the earlier reply finishes streaming in first, and queued messages then send automatically. The history still shows immediately, and the composer stays usable.
  • If a turn is still refused with agent_thread_locked (for example, another device is running the thread), the message goes back first in the queue and the queue is put on hold, instead of showing the raw lock error. The existing Send queued messages control resumes it. The server refused the turn before running it, so no work was done.
  • runConversationTurn reports the first, specific failure code. ConversationTurnError carries that code.

Verification

  • New tests:
    • tests/conversation-sdk.test.ts uses the real CopilotKitCore: a run refused with AgentThreadLockedError rejects with code agent_thread_locked. It failed before runConversationTurn kept the first report.
    • apps/mobile/test/conversation-queue.test.ts: a refused message is restored first in line, once.
  • pnpm typecheck passes. Biome reports no issues for the changed files.
  • pnpm test: 172 passed, 1 failed. The failure is Docker subprocess uses literal argv…, which fails on Windows with or without this change.
  • Web app, with AGENT_BACKEND=model and Rich Threads:
    1. Send a long browse-and-summarize request.
    2. Reload about three seconds later.
    3. Send a follow-up immediately.
    • On main: "Thread … is locked" and Retry response.
    • With this change: the follow-up shows as Up next until the earlier reply finishes, then sends and is answered. There are no lock errors in the server log.

Integration limits

Tested on web. The change is in shared chat state handling, and native layout is unchanged.

Reloading the app while a reply is still running, then sending a
message, fails with "Thread <uuid> is locked" and a Retry button. The
chat marks the conversation loaded as soon as replayed messages
arrive, while connectAgent is still attached to the running reply, so
the next turn is sent while the server still holds the thread lock.

- Keep new turns queued until connectAgent returns. It returns once
  the thread is idle, so the earlier reply finishes streaming first
  and queued messages then send automatically.
- If a turn is still refused with agent_thread_locked (for example,
  another device is running the thread), keep the message unsent and
  on hold instead of showing the raw lock error. The server rejected
  it before running it, so no work was done.
- runConversationTurn now reports the first, specific failure code;
  CopilotKit follows a lock with a generic agent_run_failed.
@davidmckayv

Copy link
Copy Markdown
Contributor

Needs a rebase before merge. It conflicts with #54 in chat.tsx and interacts with it: connect now spans the whole live run, which widens the window where #54 swallows a run error. Rebase onto main after #54 merges, document that interaction, and confirm Retry still surfaces an error, since the new global lock-error filter also hides lock errors from Retry.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants