Skip to content

Alembic multi-worker deadlock on PG first boot (migration_lock did not serialize) #1425

Description

@vybe

Symptom

During an ops /update of instance eu2 (PostgreSQL backend, #300) from v0.6.1 (0edd1dbe) to v0.7.0 (d4741d25), the first backend boot on the Alembic-on-PG path (#1183/#1186) logged one full traceback:

psycopg2.errors.DeadlockDetected: deadlock detected
sqlalchemy.exc.OperationalError: (psycopg2.errors.DeadlockDetected) deadlock detected
  File "/app/migrations/versions/0003_agent_compatibility_results.py", line 26, in upgrade
    op.execute(...)
  File "/app/db/alembic_runner.py", line 61, in upgrade_to_head

Analysis

Two uvicorn workers appear to have entered upgrade_to_head() concurrently and deadlocked inside migration 0003_agent_compatibility_results; the losing worker aborted, the winner completed. Final state is correct: alembic_version = 0010_agent_ownership_mcp_exposed (head), backend healthy, one-time occurrence.

#1263 added the cross-process migration_lock.py for exactly this multi-worker upgrade race (observed on luminous, SQLite path). On the PostgreSQL path it evidently did not fully serialize the workers — either the lock isn't engaged on the Alembic runner path or the lock scope doesn't cover command.upgrade().

Impact

Benign this time (self-resolved), but a deadlock mid-migration on a less idempotent revision could leave a partially applied migration with one worker crashed.

Repro

Multi-worker uvicorn + DATABASE_URL=postgresql://... + pending Alembic revisions (v0.6.1 → v0.7.0 jump, revisions 0001–0010 applied on one boot).

Environment

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions