Skip to content

Log connection failures at WARNING, with their cause (opal#711) - #60

Open
Stephen-Psaradellis wants to merge 1 commit into
permitio:masterfrom
Stephen-Psaradellis:connection-failures-at-warning
Open

Stephen-Psaradellis wants to merge 1 commit into
permitio:masterfrom
Stephen-Psaradellis:connection-failures-at-warning

Conversation

@Stephen-Psaradellis

Copy link
Copy Markdown

Fixes permitio/opal#711 (the examples there - opal#584, #516 and the #588 comment - all come from this library).

A failed connect in both client handlers is re-raised for the retry wrapper and logged nowhere else, and the endpoint's catch-all around the receive loop was the only record of a client that died mid-session. All of them logged at INFO, so an operator running OPAL at WARNING or ERROR saw nothing when a client could not connect.

Every connect failure now logs at WARNING and carries the exception text. The endpoint's bare except: becomes except BaseException as err - the same catch set, so nothing about which exceptions are caught or how the disconnect is handled changes - and the message names the cause. opal#584 traced a bare Client connection failed - 52764 :: <id> line to a special character in a Postgres password, with nothing in the log to go on; that line now ends with the repr of what actually happened.

Left at INFO on purpose: Client disconnected (a normal WebSocketDisconnect), Trying server, and the reader's Connection was terminated. - those are the lifecycle, not failures. Client connection failed on the endpoint is the only place a live session's death was recorded, so it is the one I promoted.

Tests: tests/log_level_test.py - a refused client connect, and an endpoint receive loop that dies with the exact RuntimeError the catch-all's own comment names. Both fail on master (the records come out at level 20) and pass with the change. The full suite is 15 passed with the change; the existing 403 test now emits its rejection at WARNING, which is the point. flake8 --select=E9,F63,F7,F82 is clean.

Disclosure. An AI system running an autonomous software project found the issue, wrote the fix and the tests, and wrote this description; I'm the account holder submitting it and I have reviewed the patch. Take it, change it or drop it as you see fit.

This fix is offered at a fixed $300, permitio's own median completed bounty on Algora. If you want to pay for it, comment /bounty $300 on the issue before merging and Algora pays it on merge. If you would rather not pay, merge it anyway - the code is yours either way.

A failed connect in both client handlers is re-raised for the retry
wrapper and logged nowhere else, and the endpoint's catch-all around the
receive loop was the only record of a client that died mid-session. All of
them logged at INFO, so an operator running OPAL at WARNING or ERROR saw
nothing when a client could not connect (permitio/opal#711, with the
examples in opal#584, #516 and #588).

Every connect failure now logs at WARNING and carries the exception text.
The endpoint's bare `except:` becomes `except BaseException as err` - the
same catch set - so the message names the cause; opal#584 traced a
"Client connection failed" line to a special character in a password with
nothing in the log to go on.

Tests: tests/log_level_test.py - a refused client connect and a receive
loop that dies with the RuntimeError the catch-all's own comment names;
both fail on the current code (level 20) and pass with the change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Serious errors showing up with log level "INFO" instead of "ERROR"

1 participant