Skip to content

fix: job wrapper status now that the watchdog kills the payload - #8528

Merged
fstagni merged 1 commit into
DIRACGrid:integrationfrom
aldbr:fix-error-msg-killed-jobs-running-out-of-time
Oct 5, 2026
Merged

fstagni merged 1 commit into
DIRACGrid:integrationfrom
aldbr:fix-error-msg-killed-jobs-running-out-of-time

Conversation

@aldbr

@aldbr aldbr commented May 5, 2026 •

Copy link
Copy Markdown
Contributor

Since #8416 the Watchdog kills the payload itself. A job stopped for exceeding a limit usually has no exit code, and postProcess reported it as "Application thread did not complete" / "No outputs generated from job execution", hiding the Watchdog's reason. executePayload then overwrote the minor status with "Exception During Execution", which is also what the monitoring and the accounting recorded.

  • Without an exit code, postProcess reports the Watchdog's reason first, then an executor error, then the generic case.
  • The CPU figures are propagated before any early return, so accounting sees them.
  • executePayload, and the PushJobAgent after a remote postProcess, keep a Failed status that postProcess already set. A traceback is still logged when it did not.
  • Unchanged on purpose: when the Watchdog killed the payload, its verdict stands even if the payload exits 0 or asks to be rescheduled on the way out. A new test pins this.
Case Before (monitoring / accounting) With this PR
Watchdog kills the payload, no exit code (time left, CPU, wall clock, disk, stalled, kill command) Failed / Application thread did not complete is sent, then executePayload queues Failed / Exception During Execution, flushed by sendFailoverRequest. Final: Exception During Execution. The Watchdog's reason is only in the log. Failed / (e.g. Job has reached the CPU limit of the queue). ApplicationError = the same reason, kept by executePayload, and in accounting.
Same, plus an executor error Application thread failed, then the Watchdog reason, then Exception During Execution The Watchdog reason, kept
Executor error, no exit code, no Watchdog Application thread failed, then Exception During Execution. Message: No outputs generated… Application thread failed, kept. Message: the executor's error
No exit code, no error Application thread did not complete, then Exception During Execution Application thread did not complete, kept
Watchdog killed it but the payload still returned an exit code (non-zero, 0, reschedule) Failed / Unchanged (0 and reschedule now pinned by a test)
Normal exits: 0, non-zero, reschedule Application Finished Successfully / With Errors / rescheduled Unchanged
postProcess never ran (e.g. Payload process could not start) Exception During Execution + traceback Unchanged
No CPU figures available __sendFinalStdOut raised KeyError, ending as Exception During Execution The final output heartbeat is skipped, and the status reflects the real outcome

Part of #8745.

BEGINRELEASENOTES
*WorkloadManagement
FIX: a payload killed by the Watchdog is reported with the Watchdog's reason (e.g. "Job has reached the CPU limit of the queue") instead of "Application thread did not complete" / "Exception During Execution", also in the accounting and for the PushJobAgent
ENDRELEASENOTES

@aldbr
aldbr requested review from atsareg and fstagni as code owners May 5, 2026 14:20
@aldbr

aldbr commented May 6, 2026

Copy link
Copy Markdown
Contributor Author
  • Tested in certification

@fstagni
fstagni marked this pull request as draft May 27, 2026 13:27
@fstagni

fstagni commented May 27, 2026

Copy link
Copy Markdown
Contributor

Converted to draft as waiting for certification test.

@aldbr
aldbr force-pushed the fix-error-msg-killed-jobs-running-out-of-time branch from 92e79cf to 4ab4f38 Compare September 30, 2026 07:23
@aldbr aldbr linked an issue Sep 30, 2026 that may be closed by this pull request
@aldbr
aldbr marked this pull request as ready for review October 1, 2026 14:18
Since the Watchdog kills the payload itself (DIRACGrid#8416), a job stopped for
exceeding a limit usually has no exit code, and postProcess reported it as
"Application thread did not complete" / "No outputs generated from job
execution", hiding the Watchdog's reason. executePayload then overwrote the
minor status with EXCEPTION_DURING_EXEC.

Without an exit code, postProcess now reports the Watchdog's reason first,
then an executor error, then the generic case. The CPU figures are
propagated before any early return so accounting sees them.
executePayload, and the PushJobAgent after a remote postProcess, keep a
FAILED status that postProcess already set.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@aldbr
aldbr force-pushed the fix-error-msg-killed-jobs-running-out-of-time branch from 4ab4f38 to 35ce0ff Compare October 1, 2026 14:18
@fstagni fstagni closed this Oct 5, 2026
@fstagni fstagni reopened this Oct 5, 2026
@fstagni
fstagni merged commit 0ceede8 into DIRACGrid:integration Oct 5, 2026
47 of 49 checks passed
@DIRACGridBot DIRACGridBot added the sweep:ignore Prevent sweeping from being ran for this PR label Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

sweep:ignore Prevent sweeping from being ran for this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Follow up] Time management after #8416

3 participants