Summary
scripts/deploy/start.sh runs docker compose up -d without rebuilding platform images. After a Dockerfile change adds a new Python dependency (e.g., the opentelemetry-* packages added in 9146d79 / #305), self-hosted developers who pull dev and re-run start.sh end up running the new source code against an old image whose Python env is missing the new package. Uvicorn's worker crashes at import with ModuleNotFoundError, but the container reports Up because docker-compose keeps respawning it. Port 8000 never binds, the UI shows "Disconnected", and every API call fails (most visibly: "Failed to create API key" on /api-keys).
Reproduction (today's incident)
- Backend image
trinity-backend:latest last built 2026-04-13.
- Pulled
dev past 9146d79 — src/backend/main.py:33 now does from opentelemetry import trace.
- Ran
./scripts/deploy/start.sh. Script printed Trinity Agent Platform - Ready!.
- Backend logs:
File "/app/main.py", line 33, in <module>
from opentelemetry import trace
ModuleNotFoundError: No module named 'opentelemetry'
docker compose ps showed trinity-backend Up 37 minutes (worker keeps respawning).
curl http://localhost:8000/api/health → connection refused.
docker exec trinity-backend pip list | grep opentelemetry → empty.
mcp-server flipped to (unhealthy) as a downstream effect.
Source is bind-mounted (Will watch for changes in these directories: ['/app']), so the new code runs against the old image's Python env.
Workaround that fixed it: docker compose build backend && docker compose up -d backend.
Root Cause
start.sh only checks for the existence of trinity-agent-base:latest before bringing services up. It has no notion of staleness for platform images (backend, frontend, mcp-server, scheduler). Any change to a service's Dockerfile that adds installed dependencies will silently break the next start.sh run on every developer's machine until they manually rebuild.
This is the same risk class that #504 ("align self-hosted deployment scripts, configs, and docs with production operating patterns") was meant to flatten, but staleness detection is narrower and worth fixing standalone.
Acceptance Criteria
Implementation Options
Pick one — listing for discussion, not prescription:
- Always pass
--build: simplest. docker compose up -d --build. Adds a few seconds to every start.sh when caches are warm. Most foolproof.
- Detect Dockerfile-newer-than-image: compare image creation time (
docker image inspect <img> --format '{{.Created}}') against the latest mtime of the corresponding Dockerfile and the dependency-bearing source files (requirements-style pip blocks). Rebuild only the affected services.
- Pin a
BUILD_VERSION in the Dockerfile and check at startup: write a marker into the image, fail fast if the source's expected marker doesn't match. Closer to how production handles it.
Option 1 is the simplest and matches how most "just works" dev scripts behave; option 2 is the lowest-overhead but more code; option 3 keeps the failure loud and early.
Related
Files
scripts/deploy/start.sh — primary change
docker/backend/Dockerfile, docker/frontend/Dockerfile, docker/mcp-server/Dockerfile, docker/scheduler/Dockerfile — possibly need a marker if Option 3 is chosen
docs/DEPLOYMENT.md — update to reflect new behavior
Summary
scripts/deploy/start.shrunsdocker compose up -dwithout rebuilding platform images. After a Dockerfile change adds a new Python dependency (e.g., theopentelemetry-*packages added in 9146d79 / #305), self-hosted developers who pulldevand re-runstart.shend up running the new source code against an old image whose Python env is missing the new package. Uvicorn's worker crashes at import withModuleNotFoundError, but the container reportsUpbecause docker-compose keeps respawning it. Port 8000 never binds, the UI shows "Disconnected", and every API call fails (most visibly: "Failed to create API key" on/api-keys).Reproduction (today's incident)
trinity-backend:latestlast built2026-04-13.devpast 9146d79 —src/backend/main.py:33now doesfrom opentelemetry import trace../scripts/deploy/start.sh. Script printedTrinity Agent Platform - Ready!.docker compose psshowedtrinity-backend Up 37 minutes(worker keeps respawning).curl http://localhost:8000/api/health→ connection refused.docker exec trinity-backend pip list | grep opentelemetry→ empty.mcp-serverflipped to(unhealthy)as a downstream effect.Source is bind-mounted (
Will watch for changes in these directories: ['/app']), so the new code runs against the old image's Python env.Workaround that fixed it:
docker compose build backend && docker compose up -d backend.Root Cause
start.shonly checks for the existence oftrinity-agent-base:latestbefore bringing services up. It has no notion of staleness for platform images (backend,frontend,mcp-server,scheduler). Any change to a service's Dockerfile that adds installed dependencies will silently break the nextstart.shrun on every developer's machine until they manually rebuild.This is the same risk class that #504 ("align self-hosted deployment scripts, configs, and docs with production operating patterns") was meant to flatten, but staleness detection is narrower and worth fixing standalone.
Acceptance Criteria
start.shbrings up a healthy backend after a Dockerfile change todocker/backend/Dockerfile(or any platform service Dockerfile) without requiring the developer to know to rundocker compose buildfirst.frontend,mcp-server, andschedulerservices — not justbackend.docs/DEPLOYMENT.mdso users know to expect the build step.Implementation Options
Pick one — listing for discussion, not prescription:
--build: simplest.docker compose up -d --build. Adds a few seconds to everystart.shwhen caches are warm. Most foolproof.docker image inspect <img> --format '{{.Created}}') against the latest mtime of the corresponding Dockerfile and the dependency-bearing source files (requirements-style pip blocks). Rebuild only the affected services.BUILD_VERSIONin the Dockerfile and check at startup: write a marker into the image, fail fast if the source's expected marker doesn't match. Closer to how production handles it.Option 1 is the simplest and matches how most "just works" dev scripts behave; option 2 is the lowest-overhead but more code; option 3 keeps the failure loud and early.
Related
feat(tracing): add OpenTelemetry distributed tracing for multi-agent calls (#305).Files
scripts/deploy/start.sh— primary changedocker/backend/Dockerfile,docker/frontend/Dockerfile,docker/mcp-server/Dockerfile,docker/scheduler/Dockerfile— possibly need a marker if Option 3 is chosendocs/DEPLOYMENT.md— update to reflect new behavior