Skip to content

fix(git-sync): the auto-sync toggle is authoritative and live (#3010) - #3017

Merged
vybe merged 4 commits into
devfrom
feature/3010-autosync-toggle
Sep 25, 2026
Merged

vybe merged 4 commits into
devfrom
feature/3010-autosync-toggle

Conversation

@dolho

@dolho dolho commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

PUT /api/agents/{name}/git/auto-sync wrote the DB row and nothing else. The container gated its heartbeat on GIT_SYNC_AUTO, read once at startup, and every recreate re-derived it as DB flag OR baked env. OFF never took effect on an agent that baked the env; ON waited for the next recreate.

  • Live on toggle. The agent loop starts whenever it can ask the platform (TRINITY_BACKEND_URL + its own TRINITY_MCP_API_KEY) and reads the owner's flag every cycle through the existing GET .../git/auto-sync. OFF skips the next cycle, ON runs it, no recreate. 404 "Git not configured" → off; platform unreachable / 5xx / auth refusal / uniform 404 → the GIT_SYNC_AUTO env (the last value the platform handed the container).
  • One writer. _apply_git_env_from_db derives GIT_SYNC_AUTO from auto_sync_enabled alone (the OR branch and its log are gone). Creation writes the flag from the same _git_auto_sync_baked predicate that bakes the env, ghosts included (the old and not config.ephemeral is what made env and DB disagree at birth).
  • Authorization aligned. Recreate still never writes the flag, so POST .../start (AuthorizedAgentByName) can't flip the owner-only PUT (OwnedAgentByName). With the OR gone, the recreate path has no way to re-arm a flag the owner cleared.
  • Reported value. Agent-side /api/git/status returns auto_sync_enabled, the value the loop is running with, from the same source GET .../git/auto-sync returns.
  • Settings → Git sync panel with both toggles (auto-sync; pause schedules while sync is failing). It uses BaseToggle, InlineError per failed save, LoadFailed/skeleton, and a "not connected" state for agents with no git binding.

⚠️ Deviation from the AC: backfill scope

The AC asks for a backfill "keyed on the creation-time predicate". The DB can't reconstruct that predicate. There's no template/fork column, and agents bound later through POST /git/initialize are also source_mode = 0 but never baked the env. A source_mode = 0 backfill would therefore arm auto-push on agents that never had it. The migration (auto_sync_enabled_backfill + Alembic 0075) sets the flag only for the slice the DB can identify: live non-source-mode ghost agents, the population the old and not config.ephemeral left env-true/DB-0.

Consequences:

  • A non-ghost agent whose flag is 0 but whose env was baked on stops auto-pushing. That covers two cases: an owner's earlier OFF, which now finally takes effect, and a rare swallowed DB write at creation. That write is now logged at error.
  • Fork-to-own ghosts (source-mode + upstream) aren't identifiable in the DB, so they aren't backfilled. This population is expected to be tiny or empty.

Changes

  • docker/base-image/agent_server/auto_sync.py: per-cycle resolve_auto_sync_enabled, should_start_loop, run_one_cycle, current_auto_sync_enabled
  • docker/base-image/agent_server/routers/git.py: auto_sync_enabled in status
  • src/backend/services/agent_service/lifecycle.py: DB-only derivation
  • src/backend/services/agent_service/crud.py: DB write from the bake predicate
  • src/backend/db/migrations.py + migrations/versions/0075_auto_sync_enabled_backfill.py: backfill, both tracks
  • src/frontend/src/components/GitSyncSettingsPanel.vue, settings/SettingsPanel.vue, stores/agents.js
  • Tests:
    • tests/unit/test_3010_autosync_toggle.py (26, new)
    • src/frontend/tests/unit/gitSyncSettingsPanel.spec.js (7, mounted)
    • test_ent109_git_env_seam.py rewritten to the new contract (recreate-after-OFF stays OFF)
    • test_2069_gitignore_at_creation.py: DB-flag matrix (7)
    • test_1484 fixture and test_1595 loop test updated for the per-cycle gate
  • Docs: git-sync-health.md, github-sync.md, architecture/backend.md, architecture/agent-lifecycle.md

Test Plan

  • Backend/agent suites around the change, run per file: all green (test_3010, ent109 ×2, 2069, 1484, agent_server_auto_sync, 1595, sync_health_service, fork_to_own, ent123, 1759, migrations, schema parity, 2742, dual ahead/behind, 69 ephemeral, alembic heads/ids)
  • Mutation: with the four source files reverted, the new tests go red:
    • test_3010: 26 errors
    • ent109: recreate-after-OFF fails
    • 2069: the ghost case fails
  • Frontend vitest run 3630/3630 (raw-colour, loading-gate and source-text ratchets included); check:tokens OK
  • lint_sys_modules, lint_root_test_placement clean; single Alembic head
  • Live on local dev: agent on a base image built from this branch, PUT the toggle, next cycle skips/runs (see comment)
  • UI checked by a human in light + dark (Settings → Git sync)

Fixes #3010

🤖 Generated with Claude Code

PUT /git/auto-sync wrote the DB row and nothing else: the container gated
on GIT_SYNC_AUTO read once at startup, and every recreate re-derived it as
DB flag OR baked env. OFF never took effect on an agent that baked the
env at creation; ON waited for the next recreate.

- agent loop: starts whenever it can ask the platform and reads the
  owner's flag every cycle via GET /git/auto-sync with its own MCP key;
  OFF skips the cycle, ON runs it, no recreate. 404 'Git not configured'
  -> off; platform unreachable / 5xx / auth refusal -> the env fallback.
  /api/git/status reports the value the loop runs with.
- one writer: _apply_git_env_from_db derives GIT_SYNC_AUTO from the DB
  flag alone (OR + log removed); creation writes the flag from the same
  _git_auto_sync_baked predicate that bakes the env, ghosts included.
- one-shot backfill (SQLite + Alembic 0075): live non-source-mode ghosts,
  the env-true/DB-0 slice the DB can identify.
- Settings -> Git sync panel with both toggles (auto-sync, pause
  schedules while sync is failing).

Fixes #3010

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dolho dolho added the ui PR touches the frontend UI — triggers Playwright e2e tests label Sep 25, 2026
…omments (#3010)

Found on a live dev agent: the first cycle waits a full interval, so
/api/git/status reported the env fallback (off) for 15 minutes while the
owner's flag said on. Resolve once before the first sleep.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@dolho

dolho commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

Live verification against the local dev instance

Setup: I built the base image from this branch as trinity-agent-base:pr3010, leaving :latest untouched, and created a throwaway agent pr3010-sync on it. The dev backend was left unchanged: the GET/PUT .../git/auto-sync endpoints this PR relies on already exist there. In the agent I set up a real repo with a bare origin. Each cycle is the container's own auto_sync.run_one_cycle → _run_auto_sync_once. The flag reads are real HTTP to http://backend:8000 with the agent's injected MCP key.

Step Owner action Cycle result
boot (container has no GIT_SYNC_AUTO) ✅ loop still starts: auto-sync loop started (interval=900s)
c1 no git config yet ✅ 404 Git not configured → skipped
c2 git config row, flag 0 ✅ skipped
c3 PUT {enabled:true} ✅ ran, pushed (remote 9bbe8b3 → 05277e9), no recreate
c4 PUT {enabled:false} ✅ skipped, remote unchanged
c5 PUT {enabled:true} ✅ ran, pushed (→ 50c0c62)
c6 backend unreachable, env unset / env true ✅ flag read failed (ConnectError); using env fallback → skipped / ran
status backend GET /git/status ✅ carries auto_sync_enabled

Found live and fixed in 1d3473dc: the first cycle waits a full interval. Until then, /api/git/status reported the env fallback (false) while the owner's flag was true. The loop now reads the flag once before its first sleep, and there is a test for it. Re-verified: with the flag on and the env unset, a restart reports true straight after boot, and it reads false after an OFF + restart.

Not covered here: the backend half of this PR (the DB-only recreate derivation, the creation write, the migration) didn't run live, because that needs the dev backend rebuilt from this branch. Unit tests and mutations cover it.

Cleanup: agent, workspace volume, git-config/sync-state rows and the :pr3010 image are removed.

🤖 Generated with Claude Code

@vybe

vybe commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

merge-train: merged dev into this branch (mechanical). It picks up #3015's sqlalchemy<2.1 cap, which is the only reason pg-migrations / schema-parity were red here. Alembic single-head check on the merged tree: 1 head (0075_auto_sync_enabled_backfill), PASS.

Heads-up: #2984, #3000, #3005 and #3023 also declare down_revision = "0074_role_readiness_rollout_seed"; once this lands they must re-parent onto 0075.

Riding the current merge train together with #3016 (validation flagged that the live toggle should not land without #3016's shared-branch push refusal).

…toggle

# Conflicts:
#	docs/memory/architecture/agent-lifecycle.md
#	tests/registry.json

@vybe vybe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

merge-train: batch validated on train/20260925-1608 (#3027, full suite green).

@vybe
vybe merged commit 675f333 into dev Sep 25, 2026
26 checks passed
webmixgamer added a commit that referenced this pull request Sep 25, 2026
…0076_operator_queue_ask_object

#3017 landed 0075_auto_sync_enabled_backfill on 0074, the same parent as
0075_operator_queue_ask_object, which made two heads: upgrade head would
then apply zero revisions (#2068). Renumber ours to 0076 and parent it on
dev's 0075. The two revisions touch disjoint tables (agent git config vs
operator_queue), so the order carries no design decision. The SQLite
migration keeps its name (operator_queue_ask_object), so a database that
already applied it does not run it again; its list entry follows dev's.

Point-in-time security reports keep the number they audited.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
vybe pushed a commit that referenced this pull request Sep 27, 2026
…only endings, wake on any ending, self-readback (Abilityai/trinity-enterprise#611, PR A of 2) (#3023)

* feat(operator-queue): ask endings — one sink, endings ledger, person-only endings, wake on any ending, self-readback (Abilityai/trinity-enterprise#611, PR A of 2)

Every way an ask ends goes through one sink, services/ask_service.py: the
operator answer, the Workspace answer, single cancel, bulk cancel, and the
poller's expiry. Each ending runs the same four steps:

- a compare-and-set that also writes the endings ledger (disposition,
  disposed_at, disposed_by person|timeout, disposed_by_email,
  disposition_reason, batch_id);
- one audit row per transition;
- one thin /ws trigger;
- the registered ending observers, which receive only the rows this call won.

What changes:

- Only a person ends an ask. respond / cancel / bulk-cancel and the Workspace
  answer refuse agent-, system- and every other non-person key with 403
  person_required (an allowlist).
- An answer after the deadline is a 409 expired, even before the sweep.
- The ent#329 wake now fires on any ending, under the same opt-in. A cancel or
  an expiry wakes the agent once per agent per event (trigger
  operator_ending); an expiry is framed as "Denied by timeout".
- An agent reads back its own ask by request_id: new
  GET /api/agents/{name}/operator-queue/{request_id}, and MCP get_my_ask.
- Every composed turn gets an "Ended asks" Execution Context line.
- Endings show with who and when in the Operations Resolved feed, a /m
  "Recently ended" strip, and the Workspace (ended asks listed for 7 days
  with a coarse who).
- Migration pair operator_queue_ask_object (SQLite) / 0075 (Alembic):
  12 nullable columns.

Review and CSO fixes, each with a failing test first and a red mutation:

- an out-of-range deadline no longer fails ingest;
- the wake's container check runs off the event loop;
- an operator's Clear All no longer hides a client's ended asks;
- the Clear-All confirmation promises a wake only to running agents;
- a system-scoped key can no longer answer an ask through the Workspace
  answer route.

Part of Abilityai/trinity-enterprise#611 (PR A of 2).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(registry): the ent#611 suite also pins the Workspace person gate, Clear-All read-through, deadline overflow and the off-loop running check

Part of Abilityai/trinity-enterprise#611 (PR A of 2).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* chore(migrations): re-parent the ent#611 revision onto dev's 0075 as 0076_operator_queue_ask_object

#3017 landed 0075_auto_sync_enabled_backfill on 0074, the same parent as
0075_operator_queue_ask_object, which made two heads: upgrade head would
then apply zero revisions (#2068). Renumber ours to 0076 and parent it on
dev's 0075. The two revisions touch disjoint tables (agent git config vs
operator_queue), so the order carries no design decision. The SQLite
migration keeps its name (operator_queue_ask_object), so a database that
already applied it does not run it again; its list entry follows dev's.

Point-in-time security reports keep the number they audited.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ui PR touches the frontend UI — triggers Playwright e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants