Repository navigation
Conversation
Drive the pipeline through a scripted transport: batching into one OTLP request, gzip and priority, retry after backoff, rejection and throttling, the write-ahead cache replayed on the next start, holds while connecting and under device pressure, the per-interval budget, the byte cap, flood guard and log floors, sessions with their own trace and attributes, typed and subscribe spans, RTC stats windows and the loss counters.
One test per row of the backend contract: every collector status code and condition, the destination and credential rules, and a mock collector driven over the real `livekit-net` HTTP stack.
59a5570 to
3205dd2
Compare
There was a problem hiding this comment.
🔍 Devin Review: 2 flags
Not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
xianshijing-lk
left a comment
There was a problem hiding this comment.
lgtm on the tests, two questions.
| telemetry.emit(TelemetryEvent::new("lk.ping")); | ||
| telemetry.set_device_state(DeviceState { | ||
| app_state: AppState::Background, | ||
| ..DeviceState::default() |
There was a problem hiding this comment.
what is the default value for Background flush ? can we do a A|B test to confirm the behaviors with / without background flush ?
There was a problem hiding this comment.
There's no switch: entering the background triggers a one-time flush (subject to the usual holds), and the periodic tick runs in both states (60 s, 120 s in the background). The test is now a paused-time A/B: foreground waits for the tick, background flushes at once. Fixed in edb3a5b.
| tokio::spawn(exporter.run()); | ||
| telemetry.emit(TelemetryEvent::new("lk.ping")); | ||
| telemetry.flush().await; | ||
| let head = seen.recv().await.expect("first request"); |
There was a problem hiding this comment.
should this seen.recv().await() have a timeout to avoid infinitely wait ?
There was a problem hiding this comment.
It's now bounded at 15 s (the 10 s export timeout plus margin) and a hang names the stuck line; the other waits that can hang got the same. Fixed in d02ae44.
Summary
Tests for #1483: the pipeline end to end, and the backend contract one test per row. Every link below points at the stage that adds the test.
Changes
telemetry.rs: the pipeline end to end through a scripted transportbackend_tests.rs: the backend contract, one test per row, including a mock collector over the reallivekit-netHTTP stacktelemetry.rstest that needsglobalarrives with it in feat(telemetry): add the process-wide pipeline and opt-out #1485Backend contract: every status code and condition → behaviour → test (29 rows)
A Room hands over its server URL and token; only a parsed
*.livekit.cloudhost over TLS getshttps://<host>/observability/client/{logs,traces}/otlp/v0withAuthorization: Bearer <token>. Every collector answer is classified by the core; behaviour marked custom is not prescribed by the OTLP/HTTP spec..livekit.cloudwith its own label, default port, no userinfo; endpoint built from that host alone; anything else collects nothinglook_alike_server_urls_never_get_the_tokenonly_livekit_cloud_hosts_get_an_ingest_urlthe_ingest_url_and_token_come_from_the_roomself_hosted_servers_get_nothing_and_nothing_is_kepthanding_over_the_same_token_again_is_freea_reconnect_with_the_same_token_retries_a_404an_expired_token_holds_uploads_until_a_fresh_one_arrivesexpis expired; only a missing one means no known expiry; a far-future one is clamped to a yearan_expired_token_holds_uploads_until_a_fresh_one_arrivesuntrusted_expiry_claims_fail_closeduploads_wait_for_a_destinationhard_holds_have_no_escape_hatcha_room_without_the_observability_grant_sends_nothinga_token_without_the_grant_is_not_consenta_refresh_that_drops_the_grant_keeps_uploading_with_the_granted_tokena_refresh_that_drops_the_grant_keeps_the_granted_token_until_it_expiresexpiredanswers_to_pre_connect_batches_are_attributed_to_their_own_projectpre_connect_records_replay_after_a_restarta_crash_while_binding_still_replays_to_the_rooms_projecta_room_switching_projects_takes_nothing_alongan_unconnected_room_never_borrows_another_rooms_projecttwo_rooms_on_two_projects_never_share_a_token_or_a_destinationownership_survives_a_room_changing_projectsrooms_never_borrow_each_others_tokensrefusals_stick_and_credentials_follow_live_roomsa_gone_rooms_pre_connect_backlog_keeps_its_credentiala_restart_with_cached_data_and_no_token_waits_for_the_same_projectbatches_events_into_one_otlp_requestpartial_successpartial_success_counts_the_refused_records_and_never_retriessuccess_and_partial_successrejected; a delay hint changes nothinga_bad_request_drops_the_batchclient_errorsdelay_hints_never_make_a_final_status_retryableoversized)payload_too_large_splits_down_to_single_recordsa_split_never_moves_a_batch_off_the_disk(added in #1485)a_split_crashed_at_every_step_loses_and_duplicates_nothingdisabled_project_goes_silent_and_purges_its_cachedisabled_by_ownerunauthorized_holds_the_batch_until_a_new_tokenrepeated_unauthorized_answers_wait_for_renewal_without_lossa_refused_token_is_never_sent_againnot_found_on_the_derived_endpoint_goes_silentnot_found_recovers_with_the_next_token_disabled_does_nota_reconnect_with_the_same_token_retries_a_404Retry-After, elseRetryInfo, else 60 s; collection continuesthrottling_honors_retry_after_then_retry_info_then_a_minutea_real_http_collector_throttles_and_recoversRetry-After; RFC 6585 §4; 60 s default customthrottling_honors_retry_after_then_retry_info_then_a_minutea_failing_server_never_costs_a_batchserver_errorsRetryInfoa_retryable_500_waits_for_its_retry_infodelay_hints_never_make_a_final_status_retryablerejectedother_server_errors_drop_the_batchRetry-After/RetryInfovaluesuntrusted_time_values_are_boundedretry_after_http_datesshutdown_never_cuts_a_server_delay_shortthrottledonly for the paused destinationa_failing_project_does_not_pause_the_otherstimeouts_are_counted_apart_from_failuresno_answer_backs_off_exponentially_with_full_jitterbackoff_doubles_with_full_jitter_up_to_a_minutean_undeclared_foreign_exception_is_a_retryable_failurean_invalid_request_is_droppedAuthorizationwhen the host or port changes (a scheme-only change on the same explicit port is not tested)redirects_to_another_host_or_port_never_carry_the_tokenclient_errorsPriority: u=7; ≤ 512 records and ≤ 1 MiB encoded, session attributes and the self-report included; the self-report is added at most once per pass and never takes the place of real records, even at a batch size of 1; a single record over the limit dropped (oversized)a_failing_self_report_never_starves_real_records(added in #1485)requests_are_gzipped_and_low_priorityencoded_requests_stay_under_the_byte_limit_for_both_signals(added in #1485)final_limits_hold_after_decoration_and_caller_strings_are_bounded(added in #1485)LK_TELEMETRY_ENDPOINT: everything there, no Cloud rules, no tokenthe_override_reaches_a_local_collector_without_cloud_rulesthe_override_takes_everything_without_a_tokenVerification
At
d02ae443, from a clean checkout (CI's test workflow runs only for PRs intomain, so these were run locally; there is no clippy job in CI):cargo fmt -- --checkcargo clippy -p livekit-telemetry --all-targets --all-features -- -D warningscargo check -p livekit-telemetry --all-targets --no-default-featureswith features[],[net],[uniffi],[net,uniffi]cargo test -p livekit-telemetry: 126 unit, 2 doc;--all-features: 129 unit, 2 doccargo doc -D warningsreports exactly what it reports on6aba1b68(private-item links,ExportError::from_response).