Roles are assigned by the task scope and the human owner, not by the agent vendor, model name, or self-assessed skill. At the start of work, use the narrowest role that can complete the request. A role change requires an explicit handoff in the task; delegation never broadens the original authority.
- Spec / Architecture Owner owns acceptance criteria, non-goals, ADR and ASR content, public contracts, and domain invariants. This role may inspect UI code but must not choose or rewrite CSS, layout, visual styling, or interaction polish unless UI work is explicitly in scope.
- Domain / API Implementer owns shared Python logic,
api/,models/,services/,data/, migrations, and contract-backed tests. This role consumes approved specs and must not redefine product scope while implementing it. - UI / Design Specialist owns
web/components, layout, accessibility, interaction behavior, and styling within existing API and TypeScript contracts. This role may inspect specs and backend contracts but must not edit BDD/spec acceptance, ADRs, API schemas, or domain logic as part of a UI task. - Independent Reviewer is read-only by default. It reports evidence-bound findings and may run only approved checks; it does not edit code, create issues, approve review, merge, or silently expand scope.
- Supervisor / Integrator assigns bounded work, checks repository state, reproduces findings, resolves handoffs, and reports what actually happened. Human approval remains required for merge, destructive actions, live writes, and materially broader scope.
Tool defaults are routing hints, not permissions: Codex commonly serves as Spec / Architecture Owner, Domain / API Implementer, Reviewer, or Supervisor; Claude commonly serves as UI / Design Specialist when assigned UI work; OpenCode defaults to Independent Reviewer. Any tool may take another role only when the task explicitly assigns it and all scope rules above still hold.
When work crosses a boundary, stop that slice at a clean handoff: state the required input, exact files or contract involved, evidence already collected, and the role that should continue. UI work uses existing component tests, fixtures, or isolated acceptance surfaces; the absence of a showcase does not authorize a UI agent to move business logic into the browser.
AI Trainer is in an active web migration. The main product development path is api/ + web/: api/ exposes FastAPI contracts over shared Python logic, web/ is the Next.js UI. Legacy Streamlit entry point is app.py — a thin fallback shell (page dispatch, sync callbacks, theme) while migration continues; Streamlit stays supported until parity, so never describe the project as already migrated.
- Intervals.icu is the recommended PRIMARY data source (
PRIMARY_ACTIVITY_SOURCE=intervals,services/intervals_icu.py); Garmin Connect is the compatible optional second source (data/garmin_client.py). Route new ingestion through the service boundary inservices/(sync, demo, acceptance orchestration) plusconfig/settings.pyfor all defaults/env lookups. models/= AI providers, coach runtime, planning, readiness/Banister metrics.utils/= shared helpers (metrics, Plotly/theme).ui/pages/,ui/components/,state/are legacy Streamlit surfaces;state/manager.pyis Streamlit-specific and must not become the contract for new product flows. Treat them as maintenance/fallback unless the task explicitly targets Streamlit.- Ignore
archived/,backups/,spikes/,debug/,examples/,research/,output/(ruff-excluded experiment areas).ai_trainer.dbat root is the local SQLite cache.
source ai_trainer_env/bin/activate # venv (Windows: Scripts\activate)
pip install -r requirements.txt -r requirements-dev.txt # base + tests
pip install -r requirements-web.txt # only for API/web runtime work
./setup_env.sh # bootstrap fresh environment
./run_web.sh # FastAPI :8000 + Next.js :3000; auto-installs missing deps; override API_PORT=/WEB_PORT= if busy
./run.sh # legacy Streamlit (:8501)
ACCEPTANCE_PORT=8510 ./run_acceptance.sh # isolated temp-DB/demo-mode surface for browser checks
python scripts/doctor_env.py check|repair --runtime|--dev # dependency diagnosticsPython lint and tests (mirrors CI .github/workflows/ci.yml exactly):
python -m ruff check . # must be green (CI job)
python -m pytest -m "not live and not debug and not e2e" tests/ # contributor-safe pass
python -m pytest tests/smoke -q # smoke subset
python -m pytest tests/test_ai_coach.py # focused module (swap filename)
python -m pytest -m e2e tests/e2e -q # Playwright web E2E, needs chromiumPytest markers (pytest.ini): smoke, live (needs credentials/network/AI runtimes), debug, e2e. Do not treat bare python -m pytest tests/ as the normal command.
Web changes MUST be verified before pushing:
npm --prefix web run lint && npm --prefix web run buildapi↔web contract: mirrored in web/lib/types.ts, scenarios in tests/contracts/registry.json, artifact tests/contracts/ts_contract.json. When either side changes, regenerate: npm --prefix web run contract:extract (CI gates freshness via contract:extract -- --check in job web-contract, then runs tests/smoke/test_contract_extractor.py + tests/smoke/test_api_call_inventory.py; inventory script: npm --prefix web run contract:inventory).
Self-hosted deployment: Dockerfile.api + docker-compose.yml + deploy/Caddyfile (Caddy auth+proxy, only Caddy exposed). Details in docs/self_hosted_deployment_execplan.md, docs/sqlite_backup_restore.md.
New product-facing behavior goes through shared Python logic + explicit API contracts in api/, consumed by web/. Streamlit changes are acceptable only for bug fixes, acceptance/admin tooling, compatibility bridges, or extracting reusable logic out of legacy UI code. Never ship a new product feature only in ui/pages/*, and never duplicate business logic between the Streamlit and api/web paths.
Before significant architecture/planning work, read docs/architecture/asr_catalog.md (ASR catalog + ADR registry — the living single source of truth for quality attributes) and docs/architecture/architecture_analysis_add3.md (tactics, ATAM-style risk heatmap, ASR→Module→Tactic map). UI/backend migration policy lives in docs/architecture/adr_0001_web_primary_ui.md; keep plans and PR scope aligned with it.
- Type hints expected on public functions (see
models/); keep docstrings in the language the file already uses (many core modules are Russian-first — match the file, don't translate). - For legacy Streamlit work prefer existing helpers in
utils/modern_ui.pyandui/components/*over inline HTML; keep data clients UI-agnostic (return structured status/errors).
Many integration suites expect populated SQLite data; favor targeted modules unless you have synced Intervals/Garmin credentials. Generate sample data via tests/add_test_data.py. Wipe the cache with tests/clean_database.py.
Conventional Commits (feat:, fix:, refactor:, chore:), subjects <72 chars. PR body lists tests run; link the issue; include screenshots/GIFs when touching any UI surface.
- Secrets only in
.env(never tracked); keys come from env — empty provider fields in the UI mean "fall back to.env", keep that behavior. - Back up/migrate
ai_trainer.dbwithscripts/sqlite_backup_restore.py(backup ... --confirm-stopped/restore ...) — a plain copy misses committed pages in the-walfile. logs/may contain personal training metrics — redact before publishing branches.
/decisions, /recovery, and the shadow /today module are hidden behind build-time flag NEXT_PUBLIC_SHOW_DEV_TOOLS=true (inlined at build — restarting the dev server after the change is required).
Canonical workflow: assign a Change Class using docs/AI_Feature_Development_Workflow.md (Full / Standard / Fast track, then SpecDD → BDD → TDD → Contract First → Self-Review → Minimal Complexity). Class A work uses docs/templates/slice_spec_review_template.md; docs-only/one-line fixes normally use Class C unless an automatic escalation trigger applies. Issue-first agent loop and GitHub automation model: docs/loop_engineering_instruction.md. Record tracked post-merge outcomes in docs/engineering_process_metrics.md.
- Evidence Discipline: causal claims must use
Observed/Inferred/Verified byand run one cheap falsifying check before naming a bug or cause; follow.agent/PLANS.mdanddocs/AI_Feature_Development_Workflow.md.
Automated reviewers re-scan the full diff on every trigger, so review loops converge only under an explicit stop-rule. docs/AI_Feature_Development_Workflow.md («Review severity», «Review budget») binds humans and bots alike:
- Issue acceptance criteria + non-goals define blockers. A finding that violates neither is a labelled suggestion (follow-up issue), never a merge gate.
- Blocking severity (P1, blocking P2) requires a reproduction or a named violated invariant; unproven what-if scenarios default to follow-up issues, not fixes on the open branch.
- Answer every finding in writing as
fixed-in <sha>/follow-up #N/disputed: <reason>— silently auto-fixing every finding feeds the next review round. - After the first consolidated full-diff round, request only scoped delta reviews (
@codex reviewexplicitly naming the changes since<sha>); the Review-budget cap of two full-diff rounds per PR is a hard limit. - Review is complete at zero open P1 with every P2 addressed or triaged in writing; merge on green CI without waiting for a "no findings" pass — chasing zero findings does not terminate.
- Never reverse a fix requested by an earlier round without fresh reproducing evidence.
OpenCode follows the Independent Reviewer role above and is never the final
authority. Follow docs/opencode_external_reviewer_runbook.md. A catalog entry
from opencode models is not availability evidence: require a real minimal
smoke response before assigning work. Read-only reviews use the plan agent,
never --auto, name the exact ref/diff and allowed test command, and exclude
.env*, logs/, ai_trainer.db, personal data, and backups/. Capture and
compare git status --short --branch before and after, stop leftover processes,
and independently reproduce or dispose every finding under the same severity
and review-budget rules above. An OpenCode full-diff audit counts as a review
round. Code-writing delegation requires explicit user authorization and an
isolated worktree; do not let two agents edit the same checkout concurrently.
For screen work, use docs/design/README.md: screen responsibilities, shared UI patterns, proportional design brief, and isolated scenario checks. This supplements the existing A/B/C workflow and role boundaries; it adds no separate approval gate. Reuse existing tokens/components, and hand off missing domain/API semantics rather than inferring them in the browser.
Complex features and significant refactors use an ExecPlan from design to implementation, per .agent/PLANS.md.
The @claude workflow (.github/workflows/claude.yml) runs in a bounded sandbox (turn budget, 30-min timeout, shared quota). Norms that keep restarts cheap:
- One milestone per @claude mention. Milestone-sized self-contained tasks (a RED→GREEN pair or one review round) finish in a single run; split larger tracks.
- Commit and push every completed RED/GREEN slice immediately. A restart then costs at most one slice. Near exhaustion, stop at a clean pushed boundary and update the progress checklist.
- Verify web changes inside the run (
lint+build) before pushing; CI is the second line. - Draft PRs open automatically for action branches (
claude/issue-N-YYYYMMDD-HHMM) viaclaude-auto-draft-pr.yml; failures report back classified viaclaude-failure-notify.yml. If quota exhausts mid-track, another agent picks up from the last pushed slice.