Skip to content

Latest commit

Β 

History

119 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

adlc-team-skills

Agent skills that give coding agents your team's context at session start, so they stop working like strangers.

The problem

Coding agents start every session knowing nothing about your team. They don't know your conventions, which patterns you deprecated, or why service X never calls service Y. Developers compensate with personal CLAUDE.md files, but those live on one machine, drift out of date, and don't transfer between teammates or tools.

What this does

#1: The agent doesn't know how your team works

team-boot runs at session start and injects an index of your team's rules, personas, and decisions β€” names and one-line descriptors, roughly a hundred tokens. Not the rules themselves. When the active task matches a rule, the agent pulls that rule's full text on demand. A task touching SQL loads the SQL rule; nothing else loads.

team-boot        β†’ auto-runs at session start (event hook), injects the index
team-discover    β†’ /team-discover for a structured match table
team-constitution β†’ interactively define your team's principles
team-repair      β†’ re-index, conflict scan, freshness check

The index lives in a git repo (team-ai-directives) β€” versioned, reviewed by PR, shared by the whole team.

#2: The agent guesses instead of asking

LLMs fill ambiguity with invention. mission-brief forces a contract first: goal, constraints, non-goals, success criteria β€” then runs specify β†’ plan β†’ tasks β†’ implement ↔ converge. When the agent gets something wrong, you fix the spec, not the code.

mission-brief "add user profile API with JWT"
mission-brief --resume    # continue an interrupted mission

#3: The maker grades its own work

The evals skills build executable evaluation suites (PromptFoo or DeepEval). Anything that can be checked by code gets a binary grader β€” plain assertions, Tier 1. LLM judges are Tier 2, reserved for what static checks can't verify. Nothing auto-merges; the pipeline ends at a PR a human reviews.

evals-init β†’ evals-specify β†’ evals-clarify β†’ evals-implement β†’ evals-validate β†’ evals-analyze

#4: Session learnings evaporate

When a session surfaces a hard-won fix, levelup-specify extracts it as a Context Directive Record (CDR) and commits it back to the team repo β€” the next session starts smarter.

levelup-init β†’ levelup-specify β†’ levelup-clarify β†’ levelup-publish

#5: Product and architecture decisions are invisible

product-* skills capture product decisions as PDR files and compile them into PRD.md. architect-* skills capture architecture decisions as ADRs (Rozanski & Woods viewpoints) and compose them into AD.md. Every decision traces from record to document to code.

product-init|specify β†’ product-clarify β†’ product-implement β†’ product-analyze
architect-init|specify β†’ architect-clarify β†’ architect-implement β†’ architect-analyze

#6: Rules pile up and rot

Base models improve; yesterday's scaffolding becomes today's context noise. team-repair --build-to-delete re-runs evals without a rule; if the model passes anyway, the rule is proposed for deletion. Rules should shrink over time, not grow.

A note on context stuffing

Long contexts measurably degrade LLM performance β€” even with perfect retrieval (arXiv:2510.05381, 13.9%–85% degradation by length alone). This repo exists because of that failure mode, not in spite of it: agents get an index by default and pull full rules only when relevant. If your instinct is "more rules in context don't work" β€” we agree. That's the design.

How it works

It starts the moment you open a session. team-boot fires via the session_start event hook and injects a lean index β€” your team's constitution, CDR index, PDR/ADR indexes, and skills registry β€” roughly a hundred tokens of names and one-line descriptors. Not the rules themselves. When the active task matches a rule, the agent pulls that rule's full text on demand. A task touching SQL loads the SQL rule; nothing else does.

On an unconfigured project, team-boot points you to /team-setup, which clones, links, or scaffolds your team-ai-directives repo. team-constitution fills in your principles.

When you ask it to build something, it doesn't jump to code. mission-brief forces a contract first β€” goal, constraints, non-goals, success criteria β€” then generates an ordered step list and walks specify β†’ plan β†’ implement ↔ converge, with gates, a circuit breaker, resume, and an audit trail. At mission start it scans installed skills directories, reads each SKILL.md frontmatter, and hands the inventory to the subagent; the model picks the skill that fits each step β€” this repo's product-specify, superpowers' brainstorming, spec-kit's commands, or your own.

For new products, product-specify captures Product Decision Records (PDRs), product-clarify refines them, and product-implement compiles them into PRD.md. The architect-* skills do the same for Architecture Decision Records (ADRs, Rozanski & Woods viewpoints) and compose them into AD.md.

The evals skills build executable evaluation suites β€” binary graders (Tier 1, plain assertions) for anything code can check, LLM judges (Tier 2) only for what static checks can't. evals-validate runs the pyramid with holdout splits, TPR/TNR, and SLA headroom. Nothing auto-merges; the pipeline ends at a PR a human reviews.

When a session surfaces a hard-won fix, levelup-specify extracts it as a Context Directive Record (CDR) β€” with a paired eval CDR β€” and commits it back to the team repo. The next session starts smarter.

Rules rot, so team-repair --build-to-delete re-runs evals without a rule; if the model passes anyway, the rule is proposed for deletion. Rules should shrink over time, not grow.

And because the skills trigger automatically from the index, you don't do anything special once they're installed. Your coding agent just has your team's context.

The Basic Workflow

  1. team-boot β€” auto-runs at session start (event hook); injects the directives index (constitution, CDR index, PDR/ADR indexes, skills registry). Full rules pulled on demand when the task matches.
  2. team-setup β€” on an unconfigured project, clones/links/scaffolds the team-ai-directives repo. team-constitution fills in your principles.
  3. mission-brief β€” before code, forces a spec contract (goal, constraints, non-goals, success criteria), then walks specify β†’ plan β†’ implement ↔ converge with gates, circuit breaker, resume, and audit trail.
  4. product-specify / product-init β€” greenfield/brownfield capture of Product Decision Records (PDRs). β†’ product-clarify refines β†’ product-implement generates PRD.md β†’ product-analyze checks consistency.
  5. architect-specify / architect-init β€” same shape for Architecture Decision Records (ADRs, Rozanski & Woods). β†’ architect-clarify β†’ architect-implement generates AD.md β†’ architect-analyze.
  6. evals-init β†’ evals-specify β†’ evals-clarify β†’ evals-implement β†’ evals-validate β€” builds executable graders (binary Tier 1, LLM judge Tier 2), holdout split, TPR/TNR + SLA validation. Pipeline ends at a human-reviewed PR.
  7. levelup-specify β€” at session end, extracts hard-won fixes as CDRs + paired eval CDRs. β†’ levelup-clarify reviews β†’ levelup-publish commits to team repo β†’ next session starts smarter.
  8. team-repair --build-to-delete β€” re-runs evals without a rule; if the model passes anyway, proposes the rule for deletion. Rules shrink over time, not grow.

The agent checks the directives index before any task. Skills trigger automatically when the active task matches β€” mandatory lifecycle, not suggestions.

Install

# Skills + slash commands + session_start events
npx adlc-skills-cli add tikalk/adlc-team-skills -a opencode

# Or plain skills (no commands/events)
npx skills add tikalk/adlc-team-skills -a claude -g

Works with any agent supporting the Agent Skills standard β€” Claude Code, Codex, OpenCode, Cursor, Copilot, and others.

adlc-skills-cli wraps npx skills add and additionally generates /name slash commands and wires session_start event hooks (via .events.json) for 9 coding agents. Skills repos without .events.json get commands only.

First run: team-boot fires at session start. On an unconfigured project it points you to /team-setup, which clones, links, or scaffolds your team-ai-directives repo. team-constitution fills in your principles.

Using agentic-sdlc-spec-kit alongside this repo? See Coexistence with Spec Kit for the conflict-free install flow.

Universal orchestration

mission-brief doesn't force a proprietary ecosystem. At mission start it scans installed skills directories, reads each SKILL.md frontmatter, and hands the inventory to the subagent β€” the model picks the skill that fits each step. Works alongside:

Source Examples
mattpocock/skills /tdd, /grill-me, /code-review
addyosmani/agent-skills Exit-criteria checklists
superpowers Workflow skills
spec-kit / agentic-sdlc-spec-kit / OpenSpec SDD command frameworks
This repo product-specify, architect-specify, evals-validate, levelup-specify
Your own Anything following the SKILL.md standard

Security

On 2026-07-27 a supply-chain worm briefly injected a malicious payload into this repo's .claude/ and .vscode/ directories via a stolen maintainer token (exposure window ~11:06–18:30 UTC). The payload only executed if you cloned the repo and opened it in VS Code or started a Claude Code session inside it; the npx skills install path never shipped or ran those files.

History was rewritten to strip the payload from all commits and tags, tokens and secrets were rotated, and branch protection now blocks the vector used. Details and remediation steps: issue #1.

Lesson for any repo: treat .vscode/tasks.json and .claude/settings.json in a clone as executable code, and disable editor auto-run tasks.

Reference

Team Directives

  • team-boot β€” session-start bootstrap; injects the directives index. Auto-triggered.
  • team-discover β€” manual re-scan; structured match table (/team-discover).
  • team-setup β€” clone, link, or scaffold a team-ai-directives repo.
  • team-constitution β€” define or amend team principles interactively.
  • team-repair β€” re-index, conflict scan, freshness, --build-to-delete.
  • team-skills β€” browse/install team skills from the directives repo.

LevelUp / CDR lifecycle

  • levelup-init β€” brownfield CDR discovery from an existing codebase.
  • levelup-specify β€” extract CDRs + paired evals from the current session.
  • levelup-clarify β€” review, accept, reject, or defer pending CDRs.
  • levelup-publish β€” compile accepted CDRs into directives + goldensets + draft PR.

Change (ChDRs)

  • change-init β€” mine git history for Change Decision Records via issue-linked commits; recovers the why behind past changes (reverts, fix chains).
  • change-clarify β€” review, accept, reject, or defer mined ChDRs (provenance gate on Decision claims).
  • change-publish β€” promote accepted ChDRs to .adlc/memory/chdr/; team-boot injects the chdr.md index at session start.

Product (PDRs)

  • product-init β€” brownfield PDR discovery. product-specify β€” greenfield creation.
  • product-clarify β€” refine and approve. product-implement β€” generate PRD.md.
  • product-analyze β€” PDR↔PRD consistency. product-roadmap β€” milestone progress.

Architecture (ADRs)

  • architect-init β€” reverse-engineer ADRs. architect-specify β€” create ADRs.
  • architect-clarify β€” refine. architect-implement β€” generate AD.md.
  • architect-analyze β€” ADR↔AD consistency.

Evals

  • evals-init β€” scaffold evals/{system}/ with security baseline.
  • evals-specify β€” extract criteria from specs / failure traces.
  • evals-clarify β€” cluster, isolate holdout, publish goldset.
  • evals-implement β€” generate graders + unit tests.
  • evals-validate β€” run evaluation pyramid, TPR/TNR + SLA headroom.
  • evals-analyze β€” route failures to rules or evaluator backlog.

Orchestration & misc

  • mission-brief β€” spec-contract pipeline with converge loop, circuit breaker, resume.
  • tech-radar-context β€” injects Tikal Tech Radar context for tech choices. Auto-triggered.
  • workspace β€” multi-repo workspace: --init creates .adlc/ structure + .gitignore, discover/link/audit child repos.

Repository layout

Skills are organized into category subdirectories under skills/:

skills/
β”œβ”€β”€ architect/             # architect-* (5 skills)
β”œβ”€β”€ product/               # product-* (6 skills) + product-templates/
β”œβ”€β”€ levelup/               # levelup-* (4 skills) + levelup-helpers.{sh,ps1}
β”œβ”€β”€ change/                # change-* (3 skills) β€” ChDRs from git history
β”œβ”€β”€ mission-brief/         # core SDD orchestrator (1 skill)
β”œβ”€β”€ evals/                 # evals-* (6 skills) + evals-templates/
β”œβ”€β”€ tech-radar/            # tech-radar-* (1 skill) + resources/radar.json
β”œβ”€β”€ workspace/             # workspace (1 skill) β€” multi-repo coordination
└── team/                  # team-* (6 skills) + team-helpers.{sh,ps1}

This places every single skill exactly 2 levels deep, fully resolving the default depth limit of the skills CLI and ensuring all skills install out of the box.

Output File Layout

All skills write to .adlc/ (project root) and the team AI directives repo.

Team Directives (inside the team AI directives repository):

  • AGENTS.md β€” agent instructions (loading order, rules, skills)
  • CDR.md β€” index of approved context contributions
  • .skills.json β€” skills manifest (schema v2.0.0)
  • .mcp.json.example β€” MCP servers config example
  • context_modules/constitution.md β€” team constitution (OKF frontmatter)
  • context_modules/{rules,personas,examples}/**/*.md β€” context modules
  • context_modules/{type}/index.md β€” progressive disclosure per concept type
  • context_modules/{type}/log.md β€” chronological change log per concept type
  • skills/{name}/SKILL.md + .skills-entry.json β€” published team skills
  • evals/{directive-id}/goldset.md + goldset.json β€” directive compliance goldensets

LevelUp (inside .adlc/ of the target project):

  • .adlc/drafts/cdr/CDR-{NNN}.md β€” proposed/discovered CDRs (including eval CDRs)
  • .adlc/drafts/cdr/cdr.md β€” auto-generated CDR index
  • .adlc/init-options.json β€” team AI directives path config

Product (inside .adlc/ and repo root):

  • .adlc/drafts/pdr/PDR-{NNN}.md β€” proposed/discovered PDRs
  • .adlc/drafts/pdr/pdr.md β€” auto-generated PDR index
  • .adlc/memory/pdr/PDR-{NNN}.md β€” accepted/completed PDRs
  • .adlc/memory/pdr/pdr.md β€” accepted PDR index
  • .adlc/product/sections/{feature-area}/{section}.md β€” PRD section build artifacts
  • .adlc/product/state.json β€” DAG execution state
  • PRD.md β€” Product Requirements Document (repo root)

Architecture (inside .adlc/ and repo root):

  • .adlc/drafts/adr/ADR-{NNN}.md β€” proposed/discovered ADRs
  • .adlc/drafts/adr/adr.md β€” auto-generated ADR index
  • .adlc/memory/adr/ADR-{NNN}.md β€” accepted ADRs
  • .adlc/memory/adr/adr.md β€” accepted ADR index
  • AD.md β€” Architecture Description (repo root)
  • .adlc/architect/ β€” per-view DAG artifacts

Missions (inside .adlc/ of the target project):

  • .adlc/workflow/workflow-config.yml β€” mission execution/supervision/budgets config
  • .adlc/workflow/.mission-state.json β€” step list, completed steps, brief, discovery results
  • .adlc/workflow/runs/<feature>/mission-log.json β€” final audit trail
  • .adlc/workflow/runs/<feature>/iterations.md β€” per-implement audit entries

Governance (inside target project and repo root):

  • .adlc/drafts/evals/EVAL-{NNN}.md β€” proposed/discovered eval criteria drafts
  • .adlc/drafts/evals/evals.md β€” draft evals index
  • .adlc/memory/evals/EVAL-{NNN}.md β€” accepted/completed eval criteria
  • .adlc/memory/evals/evals.md β€” accepted evals index
  • .adlc/memory/evals/holdout.json β€” isolated/reserved holdout test dataset
  • evals/{system}/goldset.md β€” published goldset (human-readable)
  • evals/{system}/goldset.json β€” published goldset (machine-readable)
  • evals/{system}/config.yml β€” evaluation framework configuration
  • evals/{system}/config.{js,py} β€” framework test config
  • evals/{system}/graders/check_*.py β€” generated binary Python graders / metrics
  • evals/{system}/tests/test_check_*.py β€” generated unit tests verifying grader correctness
  • evals/results/validation_report.md β€” statistical validation results report

Workspace (inside parent repo root):

  • .gitmodules β€” Git submodule registrations for child repos (created by --link)
  • .adlc/ β€” shared team context (PDRs, ADRs, CDRs); parent is the single source of truth
  • Child repos discovered at depth 1; each child's .adlc/ presence is reported (informational)
OKF Compliance

Generated context modules include Open Knowledge Format (OKF) v0.1 compliant frontmatter alongside custom fields.

OKF field Status Source
type βœ… CDR context type
title βœ… CDR title
description βœ… CDR descriptor
resource βœ… Relative path to artifact
tags βœ… Context type tag
timestamp βœ… ISO 8601 datetime

Custom fields co-exist with OKF frontmatter: id, cdr_ref, created, modified, verified, age_days, evidence.

Directory structure: context_modules/{type}/index.md (progressive disclosure), context_modules/{type}/log.md (change history), cross-links between related concepts.

Workflows

Team Directives setup:

team-setup β†’ team-constitution β†’ team-boot (auto at session start)

Product lifecycle:

Brownfield: product-init β†’ product-clarify β†’ product-implement β†’ product-analyze
Greenfield: product-specify β†’ product-clarify β†’ product-implement β†’ product-analyze
Roadmap:    product-roadmap (anytime)

Architecture lifecycle:

Brownfield: architect-init β†’ architect-clarify β†’ architect-implement β†’ architect-analyze
Greenfield: architect-specify β†’ architect-clarify β†’ architect-implement β†’ architect-analyze

LevelUp / CDR lifecycle:

Brownfield: levelup-init β†’ levelup-clarify β†’ levelup-publish β†’ team-repair
Session:    levelup-specify β†’ levelup-clarify β†’ levelup-publish β†’ team-repair
History:    change-init β†’ change-clarify β†’ change-publish (team-boot injects chdr.md)
Build to Delete: team-repair --build-to-delete β†’ levelup-clarify (review deletion CDRs)

Mission:

mission-brief "feature" β†’ review brief β†’ execute steps β†’ converge β†’ mission-log.json

Multi-repo workspace:

workspace --init β†’ create .adlc/ structure + configure .gitignore
product-specify / architect-specify β†’ create shared PDRs/ADRs in parent .adlc/
workspace --link β†’ register child repos as submodules
workspace --status β†’ audit branch, dirty, unpushed, SHA drift

Application Evaluation lifecycle:

Greenfield (Spec-Driven): evals-init β†’ evals-specify (from spec) β†’ evals-clarify β†’ evals-implement β†’ evals-validate
Brownfield (Error-Driven): evals-init β†’ evals-specify (from failures) β†’ evals-clarify β†’ evals-implement β†’ evals-validate β†’ evals-analyze

Full product β†’ architecture β†’ team:

Product:     product-specify β†’ product-clarify β†’ product-implement β†’ product-analyze
Architecture: architect-specify β†’ architect-clarify β†’ architect-implement β†’ architect-analyze
Team:        levelup-specify β†’ levelup-clarify β†’ levelup-publish β†’ team-repair
12-Factor Alignment

This repo implements the Twelve-Factor Agentic SDLC.

Factor Skills How
III β€” Mission Definition Product skills PRD/PDR lifecycle ensures product decisions are documented, reviewed, and traceable before execution
IV β€” Structured Planning Architecture skills ADRs and AD.md provide structured planning artifacts using Rozanski & Woods viewpoints
VII β€” Verification-First Evals LevelUp + Evals skills LevelUp creates directive-compliance eval CDRs; evals skills build and run application-level evaluation suites (PromptFoo/DeepEval) with binary graders, holdout splits, and statistical validation
VIII β€” Ratchet Effect LevelUp + Evals skills Each session extracts eval CDRs alongside directive CDRs; each goldset publication adds criteria that monotonically increase quality β€” evals-clarify publishes, evals-validate enforces
IX β€” Traceability Product + Architecture Every decision traces from PDR β†’ PRD β†’ feature and from ADR β†’ AD β†’ code
X β€” Context Engineering Team Directives team-boot assembles constitution, CDR index, and PDR/ADR indexes into the system prompt at session start; team-discover provides manual re-scan
XI β€” Directives as Code Team + LevelUp + Product + Architecture All directive lifecycles (CDR, PDR, ADR) live in version-controlled repos, each with draft β†’ clarify β†’ accept β†’ publish β†’ analyze stages
XII β€” Build to Delete team-repair + evals-analyze --build-to-delete runs evals without directives via LLM calls; if model passes, proposes deletion (Harness Decay); evals-analyze routes spec failures to levelup-specify (rules) and generalization failures to the evaluator backlog β€” the feedback loop that makes build-to-delete verifiable

Release Process

See RELEASE.md for the release runbook, tag naming conventions, and recovery procedures.

License

MIT β€” see LICENSE.