Task · SKT-0011

User-acceptance round: agent-run scenarios on a clean machine and a fresh Grafana Cloud stack

Description

Everything the project claims it does gets exercised by an agent, from scratch, on hardware and a stack with no prior synthkit state. The project is getting outside attention, and every existing verification runs on a machine that has been carrying synthkit state for months — so the one thing never tested is the path a new user actually takes.

Rob provisions the machine and the stack separately; this task is the scenario catalogue an agent runs against them, and the findings register it produces.

The distinguishing constraint is that a fresh stack has NO pre-existing dashboards, datasources, recording rules, or folder structure. Anything synthkit needs and does not create is a defect the current lab environment hides, and it is the defect class most likely to reach a new user first.

Scenario areas to cover, each ending in an observable assertion against the live stack rather than a local exit code:

Every scenario records a verdict and, where it fails, what a new user would have seen. Findings become tracked work; the catalogue itself is durable and re-runnable for the next release rather than a one-off script.

Do not name the stack, account or tenant in tracker text, code, docs or commit messages: this repository is public and carries a forbidden-words guard.

Acceptance Criteria

Definition of Done

Implementation Plan

Run all 46 frozen acceptance scenarios from a public clone inside a fresh Go 1.27 container using only shipped instructions, record one actionable verdict row per scenario without product fixes or identifiers, and return findings for root to track before teardown.

Implementation Notes

Scenario catalogue written 2026-08-27: e2e/acceptance/SCENARIOS.md. Eight groups, A through H, every scenario shaped as precondition -> action -> assertion -> verdict, and every assertion an observable fact about the live stack rather than a local exit code — a clean docker compose up proves the process started, not that a single series arrived.

Groups: A first deployment from nothing (the agent follows the shipped instructions LITERALLY and records every improvised step, because an undocumented step an experienced operator supplies from memory is the defect being hunted); B credential and lane coverage, all 10 lanes including the ones that silently disable themselves when unconfigured; C all 26 shipped blueprints, individually and together; D control plane and admin UI across its 33 routes; E failure modes and scenarios, asserted in the DATA not just in control-plane state; F upgrade, rollback and digest reproducibility; G generated dashboards against real synthetic output; H documentation truth.

The group most likely to pay for the whole round is G5: enumerate anything synthkit needs on a stack but does not create — recording rules, folders, datasources. A fresh stack has none of the structure the established lab environment has accumulated, and that is the defect class the current environment structurally cannot reveal.

Deliberately out of scope and stated in the file so a runner does not report them: admin-UI accessibility (descoped by explicit direction), anything requiring a write to the Kubernetes clusters (that is SKT-0012 and needs the cluster owner decision), and fixing what the round finds — a round that repairs as it goes cannot tell you how bad the fresh-start experience actually is.

BLOCKED ON: Rob provisioning the clean machine and the fresh Grafana Cloud stack. The catalogue is complete and runnable without further authoring.

Note for whoever runs it: AWS SSO was expired on the authoring machine on 2026-08-27, so anything in the catalogue touching the AWS API needs a live session first.

STACK TOPOLOGY, confirmed by Rob 2026-08-27 — orientation for whoever runs this round.

Two existing stacks with distinct roles. The lab/staff stack holds the REAL data: source of truth for live captures, signals/ provenance and the reality-corpus read-back. A second, dedicated synthkit stack is the EMISSION-TEST target: where you check what synthkit output looks like once it lands in Grafana Cloud, usually deployed there from the dedicated deployment host, which already holds the credentials (a local deploy works too).

Why this round still needs a THIRD, FRESH stack rather than reusing the emission-test one: the emission-test stack has accumulated dashboards, folders and datasource structure over months. That is exactly what hides the defect class group G5 exists to find — anything synthkit needs on a stack but does not itself create. A stack that already has the structure cannot reveal a missing-structure defect.

So the division for this round: the fresh stack is the one-off cold-start test, and the existing emission-test stack remains the everyday check for whether emission looks right. Do not substitute one for the other.

Credentials for the emission path already exist on that deployment host if a comparison against the established stack is useful mid-round.

2026-08-29 — UNBLOCKED on the stack half. A dedicated bare Grafana Cloud stack has been vended for this round. Its slug, region, and the Secrets Manager paths holding its telemetry and admin credentials are recorded on EKS-0069 in the private infrastructure tracker, deliberately not here: this repository is public and the standing rule is that no stack, account or tenant identifier appears in committed text.

It is genuinely bare, which took an extra step. Both catalogs the vending machine ships set baselineDashboards.enabled: true, which creates three starter folders and dashboards. That is exactly the pre-existing structure group G5 must not have, so the request uses the minimal catalog and patches that field to false. A stack that arrives with folders already in it cannot reveal a folder synthkit fails to create.

Consequence a runner should expect and not report as a defect: the vending claim shows Ready=False. The only unready composed resource is the per-stack ProviderConfig, which nothing references because there are no dashboards or folders to manage. The stack, its credentials and its ingest endpoints are live and usable.

The stack is temporary. It is destroyed once this round completes, through a three-stage decommission recorded on EKS-0069. Do not treat it as a durable environment or build anything on it that needs to survive.

Still needed before the round runs: the clean machine. Group A requires a host with no synthkit history, since its whole purpose is that the agent follows the shipped instructions literally and records every improvised step. The remaining decision is whether that is a throwaway VM or a clean container.

Reminder for whoever runs this: the credentials cover the required lanes, but group B also exercises the optional ones that silently disable themselves when unconfigured. Check which of those the vended stack actually supports before recording a lane as failed — an unsupported product on the stack is a different verdict from a synthkit defect.

2026-08-29, decided by Rob: the clean machine is a fresh container on his Mac. Both blockers are now cleared and this round is runnable.

Shape that satisfies group A: a container with no synthkit history, carrying Go 1.27 (installation.md says earlier toolchains are rejected), git, and the Docker CLI plus Compose with the host Docker socket mounted, since A6 exercises the documented Compose path and installation.md requires Compose 2.24.4 or later. Clone from the public GitHub URL exactly as the docs instruct rather than mounting the local checkout — a bind mount carries the working tree’s state and defeats the point.

Caveat a runner must record rather than trip over: the mounted socket means the host Docker daemon’s image cache is shared, so docker pull ghcr.io/rknightion/synthkit may resolve locally and prove nothing about a cold pull. Remove the image from the daemon before any scenario that claims to test pulling, or mark that scenario partially covered and say why.

Installing the toolchain inside the container is setup, not a finding. Anything the shipped documentation fails to tell a user to install is a finding — installation.md references just check at line 38, so whether it tells a new user how to obtain just is exactly the class of gap group A exists to catch. Record it; do not fix it.

Completed 2026-08-29 from a public clone in a disposable Go 1.27 container, following shipped instructions without product fixes. All 46 scenario IDs have recorded new-user-visible outcomes: 14 pass, 22 fail, 10 blocked. All 26 blueprints were exercised individually, but declared-signal arrival was not proved for every blueprint, so acceptance criterion 3 remains unchecked. Optional products were classified before verdicts; profiles lacked write authority and other scenarios hit unsupported-product boundaries, so criterion 4 remains unchecked. The 22 failures are tracked by SKT-0030 through SKT-0038. just check, just dump, just signal-fidelity, and Docker-backed just e2e passed for the completed wave; no generation surface changed.

2026-09-06 disposition: AC4 is case (a), proven by SKT-0039 final verification on 2026-08-31: B4-B10 each recorded live pass evidence, including profiles, Faro, Synthetic Monitoring, Fleet, generations/workflow/score, separate self-observability and a combined no-loss delivery window. AC3 is case (b): the task recorded incomplete all-blueprint arrival proof; SKT-0031 later proved all 27 non-agent identities, but the agent identity remains outside that evidence and is excluded from this run by explicit decision. Task stays Done with AC3 open, without claiming all shipped identities were verified.

Final Summary

Executed and recorded the 46-scenario fresh-container/fresh-stack round with actionable user-visible evidence. Logged 14 passes, 22 failures, and 10 blocked outcomes; converted every failure into SKT-0030 through SKT-0038. Left the two unproved acceptance criteria unchecked.

2026-09-06 reconciliation supersedes the earlier AC4 unchecked statement: later live B4-B10 optional-lane evidence proves AC4. AC3 remains open because the all-agent path was not rerun.

View the source file on GitHub