Skip to content

Fix task identifier cache recovery after database errors - #637

Open
hemesh-demi wants to merge 1 commit into
graphile:mainfrom
hemesh-demi:codex/recover-task-identifier-cache
Open

hemesh-demi wants to merge 1 commit into
graphile:mainfrom
hemesh-demi:codex/recover-task-identifier-cache

Conversation

@hemesh-demi

Copy link
Copy Markdown

If a task identifier lookup rejects after withRetries gives up, getTaskDetails caches the rejected promise. Subsequent polls with the same task list reuse that rejection, so workers can stop claiming jobs even after the database recovers.

Clear the failed cache entry so the next poll can retry. Guard both completion paths by promise identity, so an older lookup cannot clear or overwrite a newer task list's cache. Successful lookups remain cached and concurrent callers still share one pending lookup.

Adds a patch changeset and five focused tests covering 53300/ECONNRESET recovery, concurrent callers, and stale success/failure. Four fail on current main; all five pass with this change. Tests control the database boundary after connection retries; no live database is required.

Validation: focused Jest suite, source and test TypeScript checks, scoped ESLint and Prettier checks. Full database integration suite not run.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant