Skip to content

fix: end cancelled agreements reliably on-chain - #715

Open
MoonBoi9001 wants to merge 21 commits into
mainfrom
mb9/review
Open

MoonBoi9001 wants to merge 21 commits into
mainfrom
mb9/review

Conversation

@MoonBoi9001

@MoonBoi9001 MoonBoi9001 commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Dipper now ends every agreement it cancels reliably: each is marked cancelling before anything is sent, so an offer landing late withdraws itself, one accepted anyway is reopened, and a retry keeps cancelling until the chain confirms the end. Chain reads skip lagging or faulty RPC endpoints, and each end is announced once confirmed, credited to whoever made it.

Warning

Before rolling back, move every Cancelling agreement (status 9) to another status: older builds can't read it.

* fix(reassess): revoke the pending offer when cancelling an agreement

Dipper puts an offer on-chain before the indexer accepts it, but cancelled such agreements only
in its database, so the indexer could still accept and be paid with nothing tracking it. The
cancel now also goes on-chain, which revokes the offer, or ends the agreement if just accepted.

* fix(chain): report a pending offer as live after a cancel

A cancel that mined but changed nothing was counted as done when the agreement was only offered,
because the post-cancel check only knew about accepted agreements. It now also reports an offer
still waiting for the indexer, so a failed revoke is retried instead of marked cancelled.

* docs(reassess): explain how a failed revoke of an offer recovers

A failed cancel can now leave an agreement that was only offered, not just an accepted one, and
the comment on the failure log said only accepted agreements were left behind.
* fix(worker): stop an offer and a reassessment from overlapping

An offer job could check an agreement was still wanted, then send its offer after a reassessment
had already cancelled it on-chain, leaving the offer open. Offers now take the reassessment lock
shared from that check until they land, and reassessments take it exclusively.

* fix(worker): let a reassessment queue behind offers already being sent

A reassessment only retried every second, so a steady stream of offers could keep it out for as
long as the stream lasted. It now waits up to 30 s for offers in flight to land, and no new offer
starts meanwhile. Only one reassessment waits; any other still defers at once.

* fix(worker): withdraw an offer whose agreement was cancelled in flight

The chain listener cancels a replaced agreement without the reassessment lock, so its cancel can
land just before that agreement's offer and find nothing to withdraw. The offer job now re-reads
the agreement once its offer lands and withdraws the offer if dipper cancelled it meanwhile.

* fix(worker): release the offer lock before withdrawing a cancelled offer

Once an offer has landed any later cancel follows it, so holding the lock through a withdraw
only kept a waiting reassessment out for another send and receipt wait.

* fix(worker): wait for offers only after a reassessment knows it will run

A reassessment that the in-flight offer limit would defer still waited up to 30 s for offers to
land first, blocking new ones for nothing. It now checks that limit before waiting, logs a wait
that times out, and the lock test now checks both locks are released after giving up.

* fix(worker): hold offers back only while cancelling unaccepted ones

A reassessment held every offer back from before its IISA call to its last cancel, so busy
reassessments could keep offers waiting past their deadline. Offers now wait only while it cancels
an agreement whose offer may be in flight; if that wait times out, the cancel is retried later.

* fix(worker): retry withdrawing an offer of a cancelled agreement

A failed withdraw only logged a warning, and a retried offer job skipped a cancelled agreement
without checking for an offer an earlier attempt had left on-chain. Both now withdraw any offer
still stored, and a failed withdraw retries the job.
* fix(listener): cancel an agreement accepted after dipper cancelled it

An indexer could accept an offer dipper had already cancelled, and the listener ignored that, so
the agreement stayed live. The listener now queues an on-chain cancel; the job checks the chain
first and logs 1 ERROR line with a fixed event field for each agreement it ends.

* fix(worker): give the on-chain cancel job about 40 minutes of retries

The job that cancels an unwanted live agreement gave up after 2 retries, about 90 seconds, which
a short RPC outage outlasts, and giving up leaves the indexer paid. It now retries 10 times.

* fix(worker): alert on every failed cancel and run one cancel per agreement

A failed cancel of a live agreement only logged a warning, so a job that ran out of retries went
unnoticed; each failure is now an alert line. Two jobs for one agreement could both cancel it and
both alert; the second now leaves it to the first. Tests check the alert lines.

* fix(reassess): mark unaccepted agreements cancelled if the cancel fails

Left unaccepted after a failed on-chain cancel, its offer job would still send the offer and the
indexer could accept an agreement dipper wanted gone. Marked cancelled, the job withdraws any offer
instead, and the chain listener cancels the agreement if the indexer accepts one anyway.
* fix(listener): record the accept of an agreement ended after cancel

An agreement dipper had marked cancelled could still be accepted and then ended on-chain, but its
accept was never recorded, so neither its accepted nor its terminated event reached Kafka. The
listener now records both from the chain once it shows the agreement accepted and then cancelled.

* fix(listener): never announce agreements accepted before events existed

Agreements accepted before lifecycle events existed have no recorded accept, and cancelling one
would have announced it. Only an agreement dipper cancelled while its offer could still be
accepted (within an hour of the deadline) can have an accept dipper missed.

* fix(listener): stop a never-accepted cancel record hiding the real end

Cancelling a replaced agreement recorded the cancel even when it was never accepted. If the
indexer accepted it after all, that early record won over the chain's, so the terminated event
could report an end before the accept. Only an accepted agreement gets the record now.

* test(listener): prove a missed accept is announced once and only whole

Adds a database test that repeat records keep the first values and announce nothing once sent,
and a test that a failed cancel record leaves the accept unrecorded. The failure log no longer
promises a retry that only comes if the listener reads the agreement again.
Offers and cancels share 1 wallet, so marking an unaccepted agreement first means its offer
lands before the cancel and is withdrawn by it, or after, when the offer job sees the mark.
This replaces a lock that paused offers during reassessment, and failed cancels are retried.
@MoonBoi9001
MoonBoi9001 added this pull request to stack #717 October 2, 2026 17:05
@MoonBoi9001
MoonBoi9001 removed this pull request from stack #717 October 2, 2026 17:57
* feat(agreements): keep retrying cancels until the chain confirms them

Offers and cancels share 1 wallet, so marking an agreement Cancelling before its cancel goes
out lets a late offer withdraw itself. The chain listener retries the cancel until the chain
shows the end, with an ERROR after 10 failed attempts; only then is the end announced.

* refactor(worker): share the read-then-cancel step between cancel paths

The offer job's withdraw and the on-chain cancel job each read the chain and cancelled only
a live agreement, with their own copy of that logic; both now use the helper the retry uses.

* fix(worker): retry an offer job that can't recheck its agreement

After its offer lands, the offer job rereads the agreement to withdraw the offer if it was
cancelled meanwhile. A failed reread finished the job, leaving such an offer open; it retries.

* fix(listener): announce a missed accept however late it is read

Accepts dipper missed were announced only if it cancelled within 1 hour of the offer deadline,
so a lagging listener dropped them. Only agreements created before events existed are skipped.

* test(config): build every test's agreement config from one helper

6 test modules each spelled out the whole agreement config, so every new setting meant
editing 6 copies. They now start from a shared test helper and override what they need.

* docs(worker): write job counts as numerals in the cancel job's comments

* fix(listener): record no accept for a withdrawn offer being cancelled

The subgraph reports a withdrawn offer as cancelled by the payer with an accept time of 0.
Recording that as an accept announced an agreement that was never live.

* fix(listener): let the listener confirm an accepted agreement that ended

The cancel retry reads the chain ahead of the listener, so marking an ended agreement there
lost who ended it and when; it now waits for the listener, which records both from the chain.

* fix(listener): retry cancels fairly, briefly and never twice at once

Each sweep took the 100 oldest cancelling agreements, so a backlog stalled the listener and hid
newer ones, and could resend a cancel still being mined. It now takes 10, least recently
checked first, and skips any marked in the last 2 minutes.

* fix(listener): count only cancels that are mined without effect

A misconfigured or paused manager failed every agreement's cancel, using up their attempts so
none was retried once fixed. Only a cancel mined without ending the agreement now counts, and
an agreement with no stored terms hash, which can never be cancelled, is given up at once.

* fix(worker): finish an offer job whose recheck fails without resending

Retrying it sent the offer again. The chain listener's cancel retry already withdraws the offer
of an agreement left cancelling, so the job now logs the failed read and finishes.

* fix(listener): retry cancels on a timer, not by keeping polls fast

Agreements still cancelling kept the listener at its fast poll rate so their retry ran often,
and one given up on kept it there for good. The retry now runs every 5 minutes at any rate.

* fix(registry): count fees of agreements still being cancelled

An agreement dipper is cancelling is paid until the cancel lands, so it now stays in the fee
estimates the indexer selection service uses to compare indexers.

* fix(registry): cancel a replaced agreement marked expired as well

A lagging listener can mark an agreement expired that was in fact accepted. Such an agreement
couldn't be marked cancelling, so its replacement's acceptance no longer cancelled it.

* fix(listener): start orphan cancels like any other, with a retry limit

The sweep for accepted agreements left behind when a request was cancelled resent their cancel
on every run with no limit. It now marks each one cancelling, and the cancel retry finishes it.

* refactor(registry): require every registry to implement the cancel retry

The production registry trait gave the retry's 2 queries do-nothing defaults, so a registry
that forgot them would silently never retry a cancel; only the test stub keeps defaults now.

* refactor(registry): stop reading a cancel count nothing uses

The cancelling agreements query returned each row's failed cancel count, which no caller reads.

* fix(listener): never confirm a cancel the chain could not be read for

A failed read was treated like a clean check, so an offer past its deadline was marked ended
without knowing it wasn't live, as an accept the listener hadn't yet recorded would be.

* fix(listener): record no accept of an old agreement being cancelled

Agreements created before dipper announced lifecycle events are never announced, but one being
cancelled had its accept recorded from the chain, which then announced it.

* fix(listener): cancel a replaced expired agreement only if it is live

Every replaced agreement marked expired was relabelled cancelled, losing its expired event and
the record that the indexer let the offer lapse. Only one the chain shows live is cancelled.

* fix(listener): limit retries of a cancel that is mined and reverts

A mined cancel that reverted was retried every 5 minutes without limit, costing gas each time.
It now counts towards the limit; one that reverts before sending is retried, logged as an ERROR.

* fix(listener): mark an ended accepted agreement after an hour's wait

The retry left an accepted agreement that already ended for the chain listener to confirm, so
one whose end the listener never read stayed cancelling for good. After an hour it marks it.

* fix(listener): stop a cancel retry sweep after 30 seconds

Each retried cancel can wait 15 seconds to be mined, and the sweep holds up the chain listener,
so 10 slow ones stalled it for minutes. A sweep now leaves what it hasn't reached to the next.

* docs(worker): say only the chain listener queues the on-chain cancel job

A reassessment no longer queues it when its own cancel fails; the cancel retry handles that.
@MoonBoi9001
MoonBoi9001 added this pull request to stack #723 October 2, 2026 22:01
@MoonBoi9001
MoonBoi9001 removed this pull request from stack #723 October 2, 2026 22:01
@MoonBoi9001
MoonBoi9001 added this pull request to stack #724 October 2, 2026 22:04
@MoonBoi9001
MoonBoi9001 marked this pull request as ready for review October 2, 2026 22:05
@MoonBoi9001 MoonBoi9001 closed this Oct 2, 2026
@MoonBoi9001 MoonBoi9001 reopened this Oct 2, 2026
@MoonBoi9001
MoonBoi9001 removed this pull request from stack #724 October 2, 2026 22:10
* fix(selection): keep indexers dipper is cancelling out of selection

An agreement dipper is still cancelling may be live and paying its indexer, yet IISA wasn't told
to skip that indexer, so it could be picked again and paid twice. It now goes on the deployment's
declined list, which blocks a pick without counting it as part of the group.

* fix(cancel): never record an indexer's cancel as dipper's

The cancel retry marked an accepted agreement it found ended as cancelled by dipper after an hour,
even when the indexer had ended it, so the wrong end was announced. It now reads who ended it from
the contract, and leaves an end by the indexer for the chain listener to record as theirs.

* fix(cancel): retry cancels of agreements that may be paying first

The cancel retry took 10 agreements per run, oldest check first, so offers nobody had accepted could
hold up a live agreement for hours. It now takes up to 50, putting first those accepted or past their
offer deadline, and its 30 second time limit decides how many it gets through.

* fix(cancel): close an indexer's cancel the listener never recorded

An agreement the indexer ended was left for the chain listener to record, so if the listener never
did, it stayed cancelling for good and its indexer was kept out of selection. After an hour, the
cancel retry now marks it cancelled by the indexer itself, naming them as the one who ended it.

* fix(cancel): stop open offers waiting forever behind paying agreements

Agreements that may be paying an indexer were always retried first, so if enough of them kept failing,
offers an indexer could still accept were never retried. Those agreements now get an hour's head
start instead, so an offer left unchecked for over an hour still gets its turn.
* fix(chain): stop treating an expired, unaccepted offer as live

The contract keeps an offer stored after its deadline, flagging that nothing can be claimed from
it, but dipper read any stored offer as live. It then paid to cancel offers nobody could accept
any more, and lost their expired status. An offer with that flag now reads as not live.

* fix(listener): stop recording a withdrawn offer as accepted

When an offer is withdrawn before anyone accepts it, the subgraph reports it as cancelled with no
accept time, and dipper still marked it accepted, so accepted and terminated events went out for
an agreement that was never live. An offer cancelled with no accept time is no longer an accept.

* fix(cancel): time cancel retries by the chain, not the subgraph

The cancel retry decided whether an unaccepted offer's deadline had passed using the subgraph's
latest block time, which stops moving while the subgraph is down, so those offers stayed
cancelling until it recovered. It now reads the chain's own latest block time instead.

* fix(chain): never read an agreement from an RPC endpoint that is behind

Dipper read an agreement's state from whichever RPC endpoint it was using, so a lagging fallback
could report an agreement as not live after dipper had seen it go live. Each read now checks the
endpoint has reached the newest block dipper has seen, and moves to the next endpoint if not.

* fix(chain): wait for an RPC endpoint a block behind to catch up

A read refused because the endpoint was behind, or because its node lacked the block asked for,
was treated as a hard failure, so dipper gave up on that endpoint at once. Hosted endpoints are
often a block behind for a moment, so both are now retried on the same endpoint with backoff.
* refactor(cancel): share one step for marking a cancel dipper confirmed

Both the first cancel attempt and the cancel retry marked an agreement cancelled by dipper,
logged it and recorded the cancel, each with its own copy that could drift apart. They now
share one step, which records the cancel whenever its transaction is known.

* fix(cancel): read the chain before sending a cancel

Starting a cancel sent it blind, so cancelling offers that never reached the chain cost gas each,
and the reassessment waited on every receipt while holding its lock. The agreement is still marked
first, so an offer that lands later withdraws itself, but a cancel now goes out only if it is live.

* fix(cancel): finish a live cancelled agreement through the cancel retry

An agreement dipper had rejected or cancelled that the subgraph showed accepted went to a separate
job with its own 40 minute retry budget, queued again on every poll. The listener now reads the
chain and, if it is live, moves it back to cancelling so the one cancel retry ends it.
* fix(cancel): count a cancel the contract refuses before it is sent

A cancel the contract rejected while it was being prepared was never counted as a failed attempt,
so one that always failed that way was retried, with an error logged, every 5 minutes for ever.
It now counts like a cancel that reverts once mined, so it reaches the give-up limit and alert.

* fix(cancel): keep checking cancels dipper gave up on until they end

Once a cancel failed 10 times dipper never looked at the agreement again, so if it later ended
it stayed cancelling for good, its fees counted and its indexer kept out of selection. It is now
read hourly, without sending more cancels, and closed once the chain shows it ended.

* fix(cancel): slow a failing cancel to hourly instead of stopping it

Counting refusals before sending meant a paused manager used up every agreement's 10 attempts
in under an hour, and dipper then never sent their cancels again once it was unpaused. It now
alerts once at the limit and keeps retrying hourly, so those cancels resume by themselves.
* docs(config): write the test settings' zero value as a numeral

The doc comment on the shared test settings spelled the configured value 0 as a word, against the
house rule that quantities and configured values are written as numerals.

* test(registry): build 3 test mocks on the shared registry stub

Three test mocks implemented every registry method by hand, most as placeholders, so each new
registry method meant editing all of them. They now build on the shared stub and keep only the
methods their tests rely on.
@MoonBoi9001 MoonBoi9001 changed the title chore: collect the agreement-cancel fixes for review fix: end cancelled agreements reliably on-chain Oct 2, 2026
* fix(chain): ignore RPC blocks too far ahead to be real

Dipper never reads from an endpoint behind the newest block it has seen, so one faulty endpoint
reporting a far-off block refused every later read. Blocks over a week ahead, plus the time
since, are now ignored, and an ERROR fires when every endpoint is far off.

* fix(chain): pass over a lagging RPC endpoint quickly and quietly

An endpoint a block behind one dipper has seen is routine, yet each read backed off from it for
seconds with a warning every time. It is now asked once more after half a second, then the next
endpoint is tried, with nothing above debug in the logs.

* fix(chain): trust the newest block only once the chain is seen moving

An endpoint stuck far behind that answered first after a restart set the newest block, and then
every endpoint that was right looked too far ahead. It is now trusted only after 2 blocks show
the chain moving, and the ERROR fires only when every endpoint, not an outage, is far off.

* fix(chain): wait a little longer for a lagging RPC endpoint

The read just after a transaction mines often lands on a node a few blocks behind, and 1 retry
half a second later could still find it behind, failing the read in a pool of 1. It is now asked
up to twice more before the next endpoint is tried.

* fix(chain): confirm the newest block by receipt or agreeing endpoints

Two rising reads from one endpoint confirmed its block, so one moving but far behind shut out
the rest, and until then a lagging endpoint wasn't refused. The block now never goes back, and
only a receipt or 2 endpoints agreeing confirm it before far-ahead blocks are refused.
* perf(chain): read each cancelling agreement once per check

Checking an agreement dipper is cancelling read its on-chain state, then read it again to learn
whether the indexer had ended it. One read now answers both, halving the calls for every
agreement that has already ended.

* fix(cancel): credit an end to the indexer when it beat dipper's cancel

If the indexer cancelled just before dipper's cancel mined, dipper's did nothing but was still
recorded as the end, naming dipper and its transaction. The read after a cancel now also says
who ended it, and an end by the indexer is left to be recorded as theirs.

* fix(cancel): keep a mined cancel's hash when it can't be read back

When the read after a mined cancel failed, the cancel was reported as failed and its transaction
hash dropped. It is now reported as unconfirmed with its hash in the log, is not counted as a
failed attempt, and is read again on the next check.

* fix(cancel): record dipper's cancel before marking the agreement ended

An end is announced once the agreement is marked ended, and the cancel's transaction was saved
just after, so the announcement could go out without it and never be resent. The transaction is
now saved first.

* fix(cancel): count a cancel that never mines as a failed attempt

A cancel an endpoint accepted but that never mined wasn't counted, so one dropped every time was
resent every few minutes for ever and the stuck-cancel alert never fired. It now counts like a
cancel the contract refused.

* perf(cancel): skip the chain when nothing is being cancelled

The cancel retry, which runs every few minutes, read the chain's latest block before checking
whether any agreement was being cancelled. It now checks first, so a quiet dipper makes no call.

* fix(cancel): describe dropped and unconfirmed cancels accurately in logs

A cancel that never mined was logged as one that didn't end the agreement, and withdrawing an
offer whose cancel mined but couldn't be read back logged no transaction. Both now say what
happened, the second with its transaction hash.

* test(chain): read through the single agreement read in a merged test

A test brought in from the branch below still called the 2 separate chain reads this branch
replaced with 1, so it no longer compiled here.

* fix(cancel): count a dropped cancel only when its receipt was checked

A cancel counted as dropped when no receipt appeared in 15 seconds, even if every receipt
check failed, so an RPC outage could use up attempts and raise the stuck alert. It now counts
only when an endpoint answered that it had no receipt.
* fix(cancel): give the listener its hour from when an end is first seen

The chain listener gets an hour to record how an agreement ended before the cancel retry closes
it out, but the hour ran from when it was marked, so one cancelling for hours closed at once. A
new column notes when a check first finds it ended, and the hour runs from then.

* fix(cancel): cancel a reopened agreement on the next sweep

An agreement found live after dipper ended it goes back to cancelling, but then sat out the 2
minutes meant for a cancel already in flight, though none was sent, and kept paying for about 7.
Only agreements never checked since being marked wait now.

* fix(cancel): stop new cancels jumping ahead of ones that may be paying

The cancel retry takes agreements checked longest ago first, with a head start for any that may
be paying, but never-checked ones always came first, so a burst of new offers held back paying
agreements. A never-checked one now counts as checked when it was marked.

* docs(registry): describe when the cancel retry notes an end accurately

Two comments overstated things: the end time is noted whenever a check finds the agreement no
longer live, not only when dipper has no cancel to show, and an agreement can be moved back to
cancelling after a failed chain read, not only after a successful one.
* fix(liveness): end stale agreements through the cancelling status

The check for indexers that stopped serving cancelled first and marked afterwards, missing the
read before sending, the retry limit and the alert. It now marks the agreement cancelling, noted
as abandoned so it still ends that way, and the code that only it used is gone.

* docs(cancel): describe how an abandoned agreement now ends

Several comments still said a cancelled agreement always ends cancelled by dipper and that the
liveness check cancels directly. They now say an agreement dipper ends because its indexer
stopped serving it passes through cancelling and ends abandoned by the indexer.

* fix(cancel): say in the stuck-cancel alert why dipper is cancelling

The alert for a cancel that keeps failing didn't say whether dipper was ending the agreement
because its indexer stopped serving it, which an operator needs to decide what to do. It now
carries that.

* fix(liveness): leave a stale agreement that can't be cancelled alone

A stale agreement with no stored terms hash was marked cancelling and replaced, but its cancel
can never be sent, so both indexers would be paid. It is left active with an ERROR for an
operator, as before.
* fix(listener): record the accept of an agreement it reopens

When the listener found an agreement dipper had cancelled live after all and moved it back to
cancelling, it then checked the old status and skipped recording the accept, so the cancel retry
treated a paying agreement as an unaccepted offer. The accept is now recorded.

* fix(cancel): log an agreement that ended while being marked as expected

An agreement that ended, or started cancelling, between being listed and being marked was logged
as a failure, with an ERROR and a failure count in reassessment. That race is expected and needs
nothing more, so it is now logged at debug level and not counted.

* test(cancel): remove test helpers that nothing reads or writes

A list of queued cancels that was never filled made its assertion always pass, and a way to mark
an agreement already cancelled on-chain was never read, so tests using it passed for another
reason. Both are gone, along with a comment naming code that no longer exists.

* docs(listener): say plainly what reopening an agreement returns

The comment ended with "True if it was" straight after "unless the chain shows it already
ended", which read as the opposite of what it returns.
* fix(cancel): clear the old end when an agreement is found live again

An agreement reopened because the chain shows it live kept its earlier end on record, such as a
withdrawn offer, so its later real end was announced with the old transaction and time, or never
if one had gone out. That record is now cleared when the chain showed it live.

* fix(cancel): count an agreement the listener already ended as ended

If the chain listener marked an agreement ended between dipper sending its cancel and
confirming it, the confirm found nothing to update, warned that it failed, and reported the
agreement as still cancelling. It now counts it as ended and logs that at debug level.

* fix(registry): replace an end recorded before the agreement's accept

A reopen after a failed chain read keeps the end on record, and later cancels only filled blank
fields, so a withdrawn offer's end could still be announced for an agreement accepted after it.
An end recorded before the accept can't be the real one, so a later end now replaces it.
* fix(selection): skip an indexer that abandoned a deployment for a month

When an indexer stops serving an agreement, dipper ends it as abandoned and asks for a
replacement, but nothing kept that indexer out, so it could be picked straight back. It is now
left out of that deployment for the standard 30-day lookback, like one that cancelled.

* docs(registry): rewrap the declined-indexer comments to the line limit

Adding the abandoned status left 2 doc comment lines well past 100 characters.
* refactor(listener): drop the date that left old agreements unannounced

Dipper skipped announcing the accept and end of agreements created before its first release
with lifecycle events, but that only spared consumers late events for old testnet agreements.
Every agreement is now announced when the chain shows it accepted or ended.

* fix(listener): let an end recorded before the accept give way

When the listener recorded an agreement's accept and end from the chain, it recorded the end
before the accept, so an older end from before the accept, such as its offer's withdrawal, was
kept. The end is now recorded again after the accept, so the real one replaces it.

* docs(listener): stop saying old accepted agreements are never announced

A comment on the accepted-event sweep said agreements from before events existed are never
announced, which stopped being true once the date that skipped them was removed.
* fix(logging): write dipper's logs without colour codes

Dipper wrote terminal colour codes into its logs even when they went to a log store, which
split fields like the event tag so a search for them found nothing. Logs are now plain text.

* feat(alerts): post the log lines an operator must act on to Slack

Dipper's tagged warnings and errors, such as a cancel that keeps failing, were only in its logs.
With a Slack webhook in the config, each listed tag is now posted, at most once per tag every
15 minutes with a count of the rest, without ever holding up the code that logged it.

* fix(alerts): escape the characters Slack reads as markup

Slack treats &, < and > in a message as links, mentions or entities, so an error carrying an
HTML page, such as a proxy's 502, came through mangled. They are now escaped.

* fix(alerts): describe a throttle window shorter than a minute correctly

The count of held-back alerts named the window in whole minutes rounded down, so 30 seconds read
as "the last 0 minutes", and a window under a minute was only checked once a minute. Both now
follow the window as configured.

* fix(alerts): count alerts dropped in a burst into Slack's figure

Alerts dropped because the queue was full only showed up as a warning in dipper's own logs, so
the "N more" figure in Slack was too low during a burst. They are now counted by event into that
figure.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant