Description
Make synthkit easy to discover, deploy, configure, author blueprints for, and operate for Grafana-fluent sales engineers and observability architects who have never used the project. This initiative excludes telemetry-signal realism, which is owned by a concurrent campaign. It owns the user journey from clone through a focused fake workload appearing in Grafana, plus repository-local agent instructions, distributable skills, marketplace metadata, and operational safety.
Acceptance Criteria
- #1 A fresh user has one documented, focused path from clone to a running synthetic workload and can verify metrics, logs, and traces with concrete expected identities
- #2 Direct and Docker deployment defaults are safe, explicit about dry-run versus live mode, and validated through their real entry points
- #3 Blueprint creation, validation, selection, activation, staging, and restart semantics are understandable without reading Go source
- #4 Repository agent instructions and skills let Claude Code and Codex create a blueprint and guide local deployment using only the cloned repository
- #5 Every audit finding is either implemented, represented by a dependency-aware child task, or explicitly excluded as telemetry-realism work
Definition of Done
- #1 make gate (build vet test race rw-proto-check spdx-check forbidden-words)
- #2 make blueprint-schema (only if a blueprint field or construct/workload config struct changed)
- #3 DRY_RUN=true go run ./cmd/synthkit -once -dump — inventory diffed against signals/
Implementation Plan
- Run six independent read-only audits covering pickup/docs, deployment operations, blueprint/workload UX, agent instructions, skills/marketplace packaging, and realistic forward tests.
- Synthesize every accepted finding into focused child tasks with explicit dependencies and entry-point acceptance checks.
- Implement the first wave: safe first-run defaults, broken onboarding examples, canonical agent context, and skill/package correctness.
- Forward-test changed skills and user journeys in isolated environments, run one integrated CodeRabbit review for code/config changes, then the proportionate repository gate.
- Finalize each completed task with objective evidence; leave later waves To Do or Parked with precise boundaries.
Final campaign phase: complete the readiness, delivery visibility, control security, optional-product, and upgrade/rollback waves; reconcile all nineteen children; run the integrated local, published-image, exact-SHA, and standing-host gates; then close the initiative only after every child is terminal and the stable deployment is restored.
Implementation Notes
Campaign started at repository head 76ce61c on 2026-08-20. Six read-only audit lanes were dispatched under doc-0001/doc-0002. Two completed lanes already prove: the primary live path can remain in DRY_RUN, direct execution can inherit a Docker-only 0.0.0.0 control bind with no token, README control examples use port 9090 instead of 8088 and an unqualified scaling key, the default dry-run loads 25 blueprints without naming a first workload, and custom blueprint staging/validation/restart states are underexplained. Remaining audits are still active; telemetry-shape findings are fenced out.
Audit synthesis complete: 16 child tasks now own every accepted finding. Wave 1 active: .01 safe clone-to-live defaults, .02 canonical agent context/sync, .03 helper/secret safety. Dependency chain then unlocks .04 focused first workload, .05 relocatable marketplace, .06 complete skills, .07 beginner docs, .08 upload validation, .09 git/apply state, .10 readiness, .11 loss/capacity, .12 control/file security, .13 optional product lanes, .14 update/rollback, .15 local docs verification, and .16 mutation-safe APIs. No telemetry-realism task or signal-catalog change was created.
Final completeness pass added .17 for mandatory live endpoint/identity/reachability preflight and .18 for reducing always-loaded maintainer history after canonicalization. The initiative now has 18 child tasks; all accepted audit findings have a durable owner.
Wave 1 landed and pushed: implementation 9c93f5c plus tracker finalization ecad9d5. SKT-0005.01 through .03 are Done after four CodeRabbit reviews (final zero findings), make gate, and 26-blueprint dry-run inventory. Wave 2 active with disjoint owners: .04 focused blueprint selection, .05 marketplace relocation, .15 local docs validation, .16 mutation-safe APIs, and .18 slim portable agent contract. Config/control security/readiness tasks remain queued behind these file owners.
Wave 2 landed and pushed on main: implementation e98411b, task finalization 1569d34. Completed SKT-0005.04, .05, .15, .16, and .18 after one integrated CodeRabbit review (5 minor issues, all fixed), make gate, and default/focused dry-run inventories. Wave 3 started with disjoint active lanes SKT-0005.06, .07, .08, and .17; root retains integration, review, Backlog, commit, and push ownership. docs.toml remains unrelated concurrent work and is excluded from campaign staging.
Wave 3 landed in implementation commit 265ecd9 and pushed at merged head dab60de after disjoint Renovate workflow commits advanced origin/main. Completed SKT-0005.06, .07, .08, and .17 after CodeRabbit reviews 7 and 8 (eight-review cap reached), make gate, focused UI/helper/docs checks, and default plus otlp-native dry-run inventories. No live Grafana calls occurred; docs.toml remains unrelated unstaged work. Next dependency-unblocked operational queue starts with .09 and .10, then .12.
Wave 3 exact-head CI exposed a real SKT-0005.17 regression: run 32372813999 job 96437819784 rejected the e2e receiver HTTP GC_PROM_RW URL under the new HTTPS-only live contract and emitted zero signals. SKT-0005.17 is reopened for a TLS-preserving e2e repair. The eight-review CodeRabbit budget is exhausted; no review 9 will be run.
Wave 3 integration regression repaired in pushed commit 12b1aca. Local make gate and repeat make e2e passed with 664 declared metrics, 3 log sources, 3 trace services, and 3 Sigil kinds; exact-SHA CI run 32374469397 and e2e job 96443059588 passed; publish run 32374470107 succeeded for the signed/attested/scanned multi-arch image. SKT-0005.17 is Done again. SKT-0005.09 through .14 remain To Do: they are substantive code/security/operations waves requiring a fresh review gate, while the user-authorized eight CodeRabbit reviews are exhausted. Recommended resume order remains .09 and .10, then .12, followed by dependency-unlocked .11, .13, and .14.
Continuation Wave A is active for SKT-0005.09 and .10 under the superseding run goal. Fresh CodeRabbit quota and a live reference deployment are authorized. Baseline exact-head RC deployment passed a focused dry-run and landed metrics, logs, and traces with zero sink failures; it also reproduced the current readiness gap (PID Up plus event-populated status, no explicit readiness verdict). External identifiers and credentials remain outside this public tracker.
Continuation Wave A completed at exact-head b6c4ea5: SKT-0005.09 and .10 passed standing-host deployment verification and are Done. The live RC exercised delivery-aware Compose health, writable-state readiness, all configured sink lanes, fresh landed metrics, and the complete Git-source fetch/restart/loaded-versus-skipped lifecycle; temporary source artifacts were removed and the deployment returned clean and healthy. Next dependency-unblocked task is SKT-0005.12, followed by .11, .13, and .14.
Campaign reconciliation complete: all nineteen child tasks are Done. The integrated closeout passed schema regeneration with no diff, make gate, a 26-blueprint/2,644-series complete-catalog dry-run, UI, docs, skills, Compose, workflow lint, and e2e gates; exact-SHA hosted CI and the corrected stable publisher also passed. The standing host finished the candidate upgrade, stable upgrade, rollback, and final stable restoration with health, writable state, configured-sink delivery, optional-lane dispositions, and fresh metrics/logs/traces verified after each transition.
Final Summary
Completed the operational-friendliness and usability campaign across all nineteen children: safe first-run defaults, focused onboarding, blueprint lifecycle clarity, portable agent instructions and skills, trustworthy control/readiness and mutation APIs, delivery visibility, optional-product gates, and reproducible signed upgrades and rollbacks. Verified through repository gates, exact-SHA CI/publication, and a complete standing-host upgrade/rollback/restoration cycle.