Skip to content

fix(network): avoid ANR by moving ConnectivityManager calls off the main thread on Android - #2606

Open
OrLeib wants to merge 1 commit into
ionic-team:mainfrom
OrLeib:fix/network-android-anr
Open

OrLeib wants to merge 1 commit into
ionic-team:mainfrom
OrLeib:fix/network-android-anr

Conversation

@OrLeib

@OrLeib OrLeib commented Oct 3, 2026

Copy link
Copy Markdown

Description

Developed with AI assistance (Claude Code), including the emulator testing below. I reviewed the change.

NetworkPlugin.handleOnResume() and handleOnPause() run on the main thread (called from BridgeActivity.onResume/onPause via Bridge.onResume/onPause) and make synchronous binder calls into system_server:

  • handleOnResume: registerDefaultNetworkCallback (via Network.startMonitoring), then getNetworkStatus(), which calls getActiveNetwork() twice and getNetworkCapabilities().
  • handleOnPause: getNetworkStatus() (same three calls), then unregisterNetworkCallback (via Network.stopMonitoring).

That is 4 blocking IPCs on the main thread per resume and 4 per pause.
When ConnectivityService is slow to answer (busy system_server, network transitions, low-end devices), the main thread waits and the app ANRs, which is exactly the stack trace in #2563 (IConnectivityManager$Stub$Proxy.getActiveNetwork <- Network.getNetworkStatus <- NetworkPlugin.handleOnResume <- Bridge.onResume).

This PR:

  1. Runs the body of handleOnResume and handleOnPause through getBridge().execute(...), i.e. on the bridge's CapacitorPlugins HandlerThread, the same serial thread that plugin methods such as getStatus already run on.
    Because that thread is a single looper, pause and resume work stays strictly ordered (and ordered with getStatus calls), prePauseNetworkStatus is only touched from one thread, and the callback is never registered twice or unregistered before it was registered.
    Pending work posted from onPause still runs on destroy, because Bridge.onDestroy uses handlerThread.quitSafely().
  2. Reuses the already-fetched activeNetwork in Network.getNetworkStatus() instead of calling getActiveNetwork() a second time.
    This removes one binder call per status query and makes the network/capabilities pair consistent (previously the two calls could return different networks).

No public API change.
networkStatusChange events, getStatus() results and the pause/resume behavior (callback unregistered while in the background, "pre-pause vs after-pause mismatch" notification on resume) are unchanged.

Change Type

  • Fix
  • Feature
  • Refactor
  • Breaking Change
  • Documentation
  • Other (CI, chores, etc.)

Rationale / Problems Fixed

Closes #2563

ConnectivityManager.getActiveNetwork(), getNetworkCapabilities(), registerDefaultNetworkCallback() and unregisterNetworkCallback() are all synchronous IConnectivityManager binder transactions, so their latency is bounded only by system_server.
Android's ANR guidance calls out binder calls on the main thread as a common ANR cause and recommends moving them to a worker thread.
The bridge plugin thread is the natural place for this in Capacitor (plugin calls and Cordova calls already run there through Bridge.execute), so no new thread or executor is introduced.

I considered registering the callback once in load() and never unregistering it on pause, but that changes behavior (events would be delivered while the app is in the background) and still leaves the getNetworkStatus() calls in handleOnResume, so I kept the existing lifecycle and only moved the work off the main thread.

Tests or Reproductions

The plugin's Android project has no meaningful test setup (only the template ExampleUnitTest / ExampleInstrumentedTest), and StrictMode does not flag binder calls (detectAll() covers disk, network sockets, custom slow calls etc., not IPC), so I verified with a system trace on an emulator instead.

Setup: a minimal Capacitor 8.5.2 app (BridgeActivity + this plugin as a local Gradle module), Android 16 / API 36 emulator (userdebug).
The app registers a test-only subclass of NetworkPlugin that wraps super.handleOnResume() / super.handleOnPause() in Trace.beginSection(...), and a page that subscribes to networkStatusChange and calls getStatus() on load and on every visibilitychange.
I captured atrace -a <pkg> binder_driver while sending the app to the background and back three times, then counted binder_transaction events issued by the app's main thread inside those sections.
All counted transactions target the same system_server binder node (ConnectivityService).

Before (main): every lifecycle callback makes 4 binder calls on the main thread (consistent over 3 runs).

HARNESS NetworkPlugin.handleOnPause   main-thread 1.97ms  binder tx on main thread inside section: 4   (codes 0x1, 0x1, 0x10, 0x2f)
HARNESS NetworkPlugin.handleOnResume  main-thread 0.73ms  binder tx on main thread inside section: 4   (codes 0x2a, 0x1, 0x1, 0x10)
... (x3 pause/resume cycles, same counts)

After (this PR): zero binder calls on the main thread; the same calls now come from the CapacitorPlugins thread (consistent over 3 runs).

HARNESS NetworkPlugin.handleOnPause   main-thread 0.02ms  binder tx on main thread inside section: 0
HARNESS NetworkPlugin.handleOnResume  main-thread 0.01ms  binder tx on main thread inside section: 0
... (x3 pause/resume cycles, same counts)
ConnectivityService binder tx outside the sections, by thread:
   CapacitorPlugin  24   (3 x resume work + 3 x pause work + 6 x getStatus)
   ConnectivityThr   6   (network callback -> getNetworkStatus, one call fewer per event than before)

Main-thread time inside the lifecycle callbacks went from roughly 0.4 to 2.6 ms on an idle emulator (unbounded on a busy device) to roughly 10 to 400 us (just posting a Runnable).

Behavior is unchanged (identical JS event sequence before and after, same harness):

  • Foreground, adb shell cmd connectivity airplane-mode enable: networkStatusChange {"connected":false,"connectionType":"none"}.
  • Foreground, airplane mode disable: networkStatusChange events for cellular then wifi, ending in {"connected":true,"connectionType":"wifi"}.
  • Background the app, enable airplane mode, resume: Detected pre-pause and after-pause network status mismatch... is logged, networkStatusChange {"connected":false,"connectionType":"none"} fires, and getStatus() after resume returns {"connected":false,"connectionType":"none"}.
  • Background the app, disable airplane mode, resume: networkStatusChange {"connected":true,"connectionType":"wifi"} fires and getStatus() returns the same.
  • No events delivered while the app is in the background, in both versions.
  • 15 rapid background/foreground cycles followed by finishing the activity with Back: no crashes, no NetworkCallback was already registered errors, and dumpsys connectivity shows every REGISTER from the app matched by a RELEASE.

Repo checks: npm run lint (prettier-plugin-java formatting check passes; the 2 SwiftLint warnings are pre-existing in Reachability.swift), npm run verify:android (./gradlew clean build test, after npm run set-settings-gradle-for-monorepo as in CI) and npm run verify:web all pass.

Not tested: a real ANR on a physical device (it depends on system_server latency and is not reproducible on demand; the trace above shows the blocking calls are gone from the main thread instead), Android versions other than API 36, and iOS/Web (untouched).

Screenshots / Media

N/A

Platforms Affected

  • Android
  • iOS
  • Web

Notes / Comments

If the bridge plugin thread is busy with a long-running plugin call from another plugin, the network resume/pause work now waits behind it instead of blocking the UI.
That only delays re-registration of the callback slightly, and keeping it on the serial plugin thread is what guarantees pause/resume ordering, so I preferred it over a separate executor.

🤖 Generated with Claude Code

…ain thread on Android

handleOnResume and handleOnPause ran registerDefaultNetworkCallback,
unregisterNetworkCallback and getNetworkStatus (getActiveNetwork and
getNetworkCapabilities) on the main thread. These are synchronous binder
calls into system_server and can block long enough to trigger an ANR.

Run that work on the bridge plugin thread via Bridge.execute, the same
serial thread plugin methods run on, so pause/resume stay ordered with
each other and with getStatus. Also reuse the active network in
getNetworkStatus instead of querying it twice.

Closes ionic-team#2563
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Network plugin handleOnResume throws ANR

1 participant