Skip to content

fix: only treat a reachability confirmation as a recovery when the app was offline#96504

Open
adhorodyski wants to merge 7 commits into
Expensify:mainfrom
callstack-internal:fix/fake-recovery-reconnect-storms
Open

fix: only treat a reachability confirmation as a recovery when the app was offline#96504
adhorodyski wants to merge 7 commits into
Expensify:mainfrom
callstack-internal:fix/fake-recovery-reconnect-storms

Conversation

@adhorodyski

@adhorodyski adhorodyski commented Jul 20, 2026

Copy link
Copy Markdown
Contributor

Explanation of Change

Production traces show serial Ping+ReconnectApp storms (up to 23 ReconnectApp calls in 52 seconds in one session) where nothing about the network changed. Root cause: NetworkState fires reconnectApp() from raw NetInfo null→true transitions, and those get manufactured with no real connectivity change by:

  1. Boot's real emission sequence undefined→null→true slipping past the dead !== undefined guard.
  2. NetInfo.refresh() re-emitting cached state after the debug paths reset prevIsInternetReachable = null.
  3. The SHOULD_USE_STAGING_SERVER Onyx callback rebuilding the NetInfo subscription on every write, where each rebuild fires its own extra Ping.

Fix:

  1. The recovery branch fires only when getIsOffline() is true at the moment reachability is confirmed, or when a one-shot pendingReachabilityRecovery token was armed. The three debug paths (setForceOffline(false), setFailAllRequests(false), simulatePoorConnection(false)) arm the token instead of exploiting the transition loophole, so debug recovery stays Ping-verified.
  2. Deleted suppressNextReachabilityRestored, the old post-reconfigure band-aid. The gate above covers it.
  3. The SHOULD_USE_STAGING_SERVER callback rebuilds the subscription only when the effective reachability URL changed. Same-value rewrites, and toggle writes that do not change the URL (like on production, where the staging flag is forced off), no longer tear down NetInfo.

Tests: one existing test (null→true fires reconnect listener) asserted the buggy behavior and was inverted. New regression tests pin that a bare null→true and repeated fake pairs never fire a reconnect, that false→true and false→null→true fire exactly once, that force-offline debug recovery still works, and that staging toggle writes which do not change the URL do not reconfigure. All affected suites pass, including SequentialQueueTest and APITest.

Fixed Issues

$ #96634
PROPOSAL:

Tests

  1. Open the app and wait for boot to settle. Confirm no ReconnectApp fires from a reachability "recovery" on a healthy network (only the normal OpenApp/reconnect flow).
  2. Toggle Force offline in the TestToolMenu on, then off. Confirm exactly one ReconnectApp fires after turning it off.
  3. Turn airplane mode on, wait for the offline indicator, then turn it off. Confirm one ReconnectApp fires after connectivity returns.
  • Verify that no errors appear in the JS console

Offline tests

Same as steps 2 and 3 above. Both exercise the offline→online recovery path.

QA Steps

Same as tests.

  • Verify that no errors appear in the JS console

PR Author Checklist

  • I linked the correct issue in the ### Fixed Issues section above
  • I wrote clear testing steps that cover the changes made in this PR
    • I added steps for local testing in the Tests section
    • I added steps for the expected offline behavior in the Offline steps section
    • I added steps for Staging and/or Production testing in the QA steps section
    • I added steps to cover failure scenarios (i.e. verify an input displays the correct error message if the entered data is not correct)
    • I turned off my network connection and tested it while offline to ensure it matches the expected behavior (i.e. verify the default avatar icon is displayed if app is offline)
    • I tested this PR with a High Traffic account against the staging or production API to ensure there are no regressions (e.g. long loading states that impact usability).
  • I included screenshots or videos for tests on all platforms
  • I ran the tests on all platforms & verified they passed on:
    • Android: Native
    • Android: mWeb Chrome
    • iOS: Native
    • iOS: mWeb Safari
    • MacOS: Chrome / Safari
  • I verified there are no console errors (if there's a console error not related to the PR, report it or open an issue for it to be fixed)
  • I followed proper code patterns (see Reviewing the code)
    • I verified that comments were added to code that is not self explanatory
    • I verified that any new or modified comments were clear, correct English, and explained "why" the code was doing something instead of only explaining "what" the code was doing.
    • I verified any copy / text that was added to the app is grammatically correct in English. It adheres to proper capitalization guidelines (note: only the first word of header/labels should be capitalized), and is either coming verbatim from figma or has been approved by marketing (in order to get marketing approval, ask the Bug Zero team member to add the Waiting for copy label to the issue)
  • If a new code pattern is added I verified it was agreed to be used by multiple Expensify engineers
  • I followed the guidelines as stated in the Review Guidelines
  • I tested other components that can be impacted by my changes (i.e. if the PR modifies a shared library or component like Avatar, I verified the components using Avatar are working as expected)
  • If a new CSS style is added I verified that:
    • A similar style doesn't already exist
    • The style can't be created with an existing StyleUtils function (i.e. StyleUtils.getBackgroundAndBorderStyle(theme.componentBG))
  • If new assets were added or existing ones were modified, I verified that:
    • The assets are optimized and compressed (for SVG files, run npm run compress-svg)
    • The assets load correctly across all supported platforms.
  • If the PR modifies code that runs when editing or sending messages, I tested and verified there is no unexpected behavior for all supported markdown - URLs, single line code, code blocks, quotes, headings, bold, strikethrough, and italic.
  • If the PR modifies a generic component, I tested and verified that those changes do not break usages of that component in the rest of the App (i.e. if a shared library or component like Avatar is modified, I verified that Avatar is working as expected in all cases)
  • If the PR modifies a component related to any of the existing Storybook stories, I tested and verified all stories for that component are still working as expected.
  • If the PR modifies a component or page that can be accessed by a direct deeplink, I verified that the code functions as expected when the deeplink is used - from a logged in and logged out account.
  • If the PR modifies the UI (e.g. new buttons, new UI components, changing the padding/spacing/sizing, moving components, etc) or modifies the form input styles:
    • I verified that all the inputs inside a form are aligned with each other.
    • I added Design label and/or tagged @Expensify/design so the design team can review the changes.
  • I added unit tests for any new feature or bug fix in this PR to help automatically prevent regressions in this user flow.
  • If the main branch was merged into this PR after a review, I tested again and verified the outcome was still expected according to the Test steps.

Screenshots/Videos

Android: Native

Android: mWeb Chrome

iOS: Native

iOS: mWeb Safari

MacOS: Chrome / Safari

@codecov

codecov Bot commented Jul 20, 2026

Copy link
Copy Markdown

Codecov Report

✅ Changes either increased or maintained existing code coverage, great job!

Files with missing lines Coverage Δ
src/libs/NetworkState.ts 93.44% <100.00%> (+0.07%) ⬆️
... and 99 files with indirect coverage changes

adhorodyski and others added 2 commits July 20, 2026 19:15
… the effective URL

Extract the repeated prev-reset/token/refresh sequence into armReachabilityRecovery(),
and compare the effective reachability URL instead of mirroring the raw staging toggle —
the raw value can change without changing the URL (and vice versa on staging, where null
falls back to the environment default).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@adhorodyski

Copy link
Copy Markdown
Contributor Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 588423ac34

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread src/libs/NetworkState.ts
adhorodyski and others added 4 commits July 21, 2026 09:53
@adhorodyski
adhorodyski marked this pull request as ready for review July 21, 2026 11:41
@adhorodyski
adhorodyski requested review from a team as code owners July 21, 2026 11:41
@melvin-bot
melvin-bot Bot requested review from JmillsExpensify and mkhutornyi and removed request for a team July 21, 2026 11:41
@melvin-bot

melvin-bot Bot commented Jul 21, 2026

Copy link
Copy Markdown

@mkhutornyi Please copy/paste the Reviewer Checklist from here into a new comment on this PR and complete it. If you have the K2 extension, you can simply click: [this button]

@mkhutornyi

mkhutornyi commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

Reviewer Checklist

  • I have verified the author checklist is complete (all boxes are checked off).
  • I verified the correct issue is linked in the ### Fixed Issues section above
  • I verified testing steps are clear and they cover the changes made in this PR
    • I verified the steps for local testing are in the Tests section
    • I verified the steps for Staging and/or Production testing are in the QA steps section
    • I verified the steps cover any possible failure scenarios (i.e. verify an input displays the correct error message if the entered data is not correct)
    • I turned off my network connection and tested it while offline to ensure it matches the expected behavior (i.e. verify the default avatar icon is displayed if app is offline)
  • I checked that screenshots or videos are included for tests on all platforms
  • I included screenshots or videos for tests on all platforms
  • I verified that the composer does not automatically focus or open the keyboard on mobile unless explicitly intended. This includes checking that returning the app from the background does not unexpectedly open the keyboard.
  • I verified tests pass on all platforms & I tested again on:
    • Android: HybridApp
    • Android: mWeb Chrome
    • iOS: HybridApp
    • iOS: mWeb Safari
    • MacOS: Chrome / Safari
  • If there are any errors in the console that are unrelated to this PR, I either fixed them (preferred) or linked to where I reported them in Slack
  • I verified proper code patterns were followed (see Reviewing the code)
    • I verified that any callback methods that were added or modified are named for what the method does and never what callback they handle (i.e. toggleReport and not onIconClick).
    • I verified that comments were added to code that is not self explanatory
    • I verified that any new or modified comments were clear, correct English, and explained "why" the code was doing something instead of only explaining "what" the code was doing.
    • I verified any copy / text that was added to the app is grammatically correct in English. It adheres to proper capitalization guidelines (note: only the first word of header/labels should be capitalized), and is either coming verbatim from figma or has been approved by marketing (in order to get marketing approval, ask the Bug Zero team member to add the Waiting for copy label to the issue)
  • If a new code pattern is added I verified it was agreed to be used by multiple Expensify engineers
  • I verified that this PR follows the guidelines as stated in the Review Guidelines
  • I verified other components that can be impacted by these changes have been tested, and I retested again (i.e. if the PR modifies a shared library or component like Avatar, I verified the components using Avatar have been tested & I retested again)
  • If a new component is created I verified that:
    • A similar component doesn't exist in the codebase
    • All props are defined accurately and each prop has a /** comment above it */
    • The file is named correctly
    • The component has a clear name that is non-ambiguous and the purpose of the component can be inferred from the name alone
    • The only data being stored in the state is data necessary for rendering and nothing else
    • For Class Components, any internal methods passed to components event handlers are bound to this properly so there are no scoping issues (i.e. for onClick={this.submit} the method this.submit should be bound to this in the constructor)
    • Any internal methods bound to this are necessary to be bound (i.e. avoid this.submit = this.submit.bind(this); if this.submit is never passed to a component event handler like onClick)
    • All JSX used for rendering exists in the render method
    • The component has the minimum amount of code necessary for its purpose, and it is broken down into smaller components in order to separate concerns and functions
  • If any new file was added I verified that:
    • The file has a description of what it does and/or why is needed at the top of the file if the code is not self explanatory
  • If a new CSS style is added I verified that:
    • A similar style doesn't already exist
    • The style can't be created with an existing StyleUtils function (i.e. StyleUtils.getBackgroundAndBorderStyle(theme.componentBG)
  • If the PR modifies code that runs when editing or sending messages, I tested and verified there is no unexpected behavior for all supported markdown - URLs, single line code, code blocks, quotes, headings, bold, strikethrough, and italic.
  • If the PR modifies a generic component, I tested and verified that those changes do not break usages of that component in the rest of the App (i.e. if a shared library or component like Avatar is modified, I verified that Avatar is working as expected in all cases)
  • If the PR modifies a component related to any of the existing Storybook stories, I tested and verified all stories for that component are still working as expected.
  • If the PR modifies a component or page that can be accessed by a direct deeplink, I verified that the code functions as expected when the deeplink is used - from a logged in and logged out account.
  • If the PR modifies the UI (e.g. new buttons, new UI components, changing the padding/spacing/sizing, moving components, etc) or modifies the form input styles:
    • I verified that all the inputs inside a form are aligned with each other.
    • I added Design label and/or tagged @Expensify/design so the design team can review the changes.
  • For any bug fix or new feature in this PR, I verified that sufficient unit tests are included to prevent regressions in this flow.
  • If the main branch was merged into this PR after a review, I tested again and verified the outcome was still expected according to the Test steps.
  • I have checked off every checkbox in the PR reviewer checklist, including those that don't apply to this PR.

Screenshots/Videos

Android: HybridApp
Android: mWeb Chrome
iOS: HybridApp
ios.mov
iOS: mWeb Safari
MacOS: Chrome / Safari
web.mov
web2.mov

@mkhutornyi

Copy link
Copy Markdown
Contributor

Suggested QA Steps:

Setup (how to observe reconnects): Open the app in Chrome, open DevTools → Network tab, and type ReconnectApp into the filter box. Leave it open during each test — every ReconnectApp request will appear as a row, so you can count them. (Alternatively filter the Console for reachability.)

1. Healthy boot — no phantom reconnects

  1. Make sure you're on a stable internet connection.
  2. Close and reopen the app (or hard-refresh) and log in.
  3. Wait ~60 seconds for the app to settle.
    • Expected: You see the normal one-time OpenApp/ReconnectApp from sign-in, and then nothing. There should be no repeating burst of ReconnectApp calls while the network is healthy and untouched.
    • Fail: A stream of repeated ReconnectApp calls with no network change.

2. Force Offline toggle → exactly one reconnect

  1. Open the Test Tools menu and turn Force offline ON.
    • ✅ The offline indicator appears; no ReconnectApp fires.
  2. Turn Force offline OFF.
    • Expected: The offline indicator disappears and exactly one ReconnectApp fires. Not zero (data must resync), not several.

3. Real connectivity loss → one reconnect on return

  1. Turn on airplane mode (mobile) or disconnect WiFi/unplug ethernet (web).
  2. Wait for the offline indicator to appear (a few seconds).
  3. Turn airplane mode off / reconnect.
    • Expected: Once connectivity returns, the offline indicator clears and exactly one ReconnectApp fires. Any pending changes you made offline sync through.

4. Rapid offline/online cycling doesn't storm

  1. Toggle Force offline ON→OFF→ON→OFF a few times in a row.
    • Expected: Each OFF produces one ReconnectApp (so N toggles → N reconnects), with no extra pile-up of calls between toggles.

5. Staging-server toggle doesn't trigger a storm (staging/dev builds only)

  1. In Test Tools, flip the Use Staging Server toggle a couple of times.
    • Expected: No burst of ReconnectApp/Ping calls purely from flipping the toggle. On a production build (where the toggle has no effect on the API URL) it should cause no reconnect at all.

In every test:

  • ✅ Data stays in sync after each recovery (make an edit while offline, confirm it lands after reconnect).

@melvin-bot
melvin-bot Bot requested a review from mountiny July 21, 2026 13:47
@melvin-bot

melvin-bot Bot commented Jul 21, 2026

Copy link
Copy Markdown

I can't run change requests for your access level. I can investigate, or file an issue in Expensify/Expensify instead.


Failed after 0s · 0 tools used · view debug log

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 954a4140c0

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread src/libs/NetworkState.ts
// or a debug path asked for one. Everything else (boot, re-subscription, refresh() re-emits)
// is a re-read of an unchanged network and must not fire reconnectApp.
if (!shouldForceOffline && state.isInternetReachable === true && prevIsInternetReachable !== true) {
if (getIsOffline() || pendingReachabilityRecovery) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve pre-event offline state for radio recovery

When the offline period was caused only by noRadioActive and NetInfo never reported isInternetReachable=false (for example isConnected=false with reachability still null, followed by isConnected=true/isInternetReachable=true), the listener calls setHasRadio(true) before this check, so getIsOffline() is already false and onReachabilityRestored() is skipped. Reconnect.ts only calls reconnectApp() from onReachabilityConfirmed, so this radio-only recovery can flush the queue but miss the data resync that fetches skipped Onyx updates; capture the pre-NetInfo-event offline state or explicitly treat radio recovery as something to recover from.

Useful? React with 👍 / 👎.

@mountiny mountiny left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants