Skip to content

bug: start.sh runs stale platform images when Dockerfiles change (silent ModuleNotFoundError on cold start) #557

Description

@webmixgamer

Summary

scripts/deploy/start.sh runs docker compose up -d without rebuilding platform images. After a Dockerfile change adds a new Python dependency (e.g., the opentelemetry-* packages added in 9146d79 / #305), self-hosted developers who pull dev and re-run start.sh end up running the new source code against an old image whose Python env is missing the new package. Uvicorn's worker crashes at import with ModuleNotFoundError, but the container reports Up because docker-compose keeps respawning it. Port 8000 never binds, the UI shows "Disconnected", and every API call fails (most visibly: "Failed to create API key" on /api-keys).

Reproduction (today's incident)

  1. Backend image trinity-backend:latest last built 2026-04-13.
  2. Pulled dev past 9146d79 — src/backend/main.py:33 now does from opentelemetry import trace.
  3. Ran ./scripts/deploy/start.sh. Script printed Trinity Agent Platform - Ready!.
  4. Backend logs:
    File "/app/main.py", line 33, in <module>
        from opentelemetry import trace
    ModuleNotFoundError: No module named 'opentelemetry'
    
  5. docker compose ps showed trinity-backend Up 37 minutes (worker keeps respawning).
  6. curl http://localhost:8000/api/health → connection refused.
  7. docker exec trinity-backend pip list | grep opentelemetry → empty.
  8. mcp-server flipped to (unhealthy) as a downstream effect.

Source is bind-mounted (Will watch for changes in these directories: ['/app']), so the new code runs against the old image's Python env.

Workaround that fixed it: docker compose build backend && docker compose up -d backend.

Root Cause

start.sh only checks for the existence of trinity-agent-base:latest before bringing services up. It has no notion of staleness for platform images (backend, frontend, mcp-server, scheduler). Any change to a service's Dockerfile that adds installed dependencies will silently break the next start.sh run on every developer's machine until they manually rebuild.

This is the same risk class that #504 ("align self-hosted deployment scripts, configs, and docs with production operating patterns") was meant to flatten, but staleness detection is narrower and worth fixing standalone.

Acceptance Criteria

  • start.sh brings up a healthy backend after a Dockerfile change to docker/backend/Dockerfile (or any platform service Dockerfile) without requiring the developer to know to run docker compose build first.
  • No regression on the happy path — fresh clones with up-to-date images don't pay a noticeable rebuild penalty (rebuild is skipped when source matches image).
  • If a rebuild is needed, the script prints what it's rebuilding and why, so the user understands the delay.
  • Same protection applies to frontend, mcp-server, and scheduler services — not just backend.
  • If the chosen approach is "always rebuild", that decision is documented in docs/DEPLOYMENT.md so users know to expect the build step.

Implementation Options

Pick one — listing for discussion, not prescription:

  1. Always pass --build: simplest. docker compose up -d --build. Adds a few seconds to every start.sh when caches are warm. Most foolproof.
  2. Detect Dockerfile-newer-than-image: compare image creation time (docker image inspect <img> --format '{{.Created}}') against the latest mtime of the corresponding Dockerfile and the dependency-bearing source files (requirements-style pip blocks). Rebuild only the affected services.
  3. Pin a BUILD_VERSION in the Dockerfile and check at startup: write a marker into the image, fail fast if the source's expected marker doesn't match. Closer to how production handles it.

Option 1 is the simplest and matches how most "just works" dev scripts behave; option 2 is the lowest-overhead but more code; option 3 keeps the failure loud and early.

Related

Files

  • scripts/deploy/start.sh — primary change
  • docker/backend/Dockerfile, docker/frontend/Dockerfile, docker/mcp-server/Dockerfile, docker/scheduler/Dockerfile — possibly need a marker if Option 3 is chosen
  • docs/DEPLOYMENT.md — update to reflect new behavior

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions