Milestones

tailscale2otel

Progress across published tasks.

user-friendliness

13 of 13 completed

Wave 1: bugs and guardrail foundations

12 of 12 completed

Wave 5: HA coordination and validation follow-ups

7 of 7 completed

v5 big bang major

1 of 1 completed

Wave 2: Loki state-snapshot family

8 of 8 completed

Wave 3: API surface expansion

9 of 9 completed

Wave 4: collector and ingestion resilience

11 of 11 completed

Wave 5: operator surface and cost controls

11 of 11 completed

Wave 6: config ergonomics and performance

6 of 6 completed

Wave 7: deploy and docs polish

5 of 5 completed

Decision-gated

2 of 2 completed

Wave 4: live validation and review follow-ups

9 of 9 completed

Unassigned

58 of 60 completed

TSO-0001Upgrade Go toolchain to 1.27UnspecifiedTSO-0002Remediate Codex Security scan b3c6de8ehighTSO-0002.01Harden ingress and local HTTP boundarieshighTSO-0002.02Bound retained state and protect local persistencehighTSO-0002.03Harden outbound credentials transports and TLShighTSO-0002.04Close Helm credential and workload identity gapshighTSO-0003Cut CI job fan-out per PR to drain the Actions queueUnspecifiedTSO-0004prometheus.max_requests_in_flight: 0 now crash-loops every config copied from the old exampleUnspecifiedTSO-0005Persistent flow store refuses every pre-hardening database with no way to migrate itUnspecifiedTSO-0006Migrate to OpenTelemetry v1.46.0 / log v0.22.0 after the log attribute API removalUnspecifiedTSO-0020Rotate credentials exposed in an agent transcripthighneeds-triagesecurityTSO-0021Update check fails permanently: the 64 KiB body cap truncates the GitHub latest-release responsehighTSO-0022Fix live Grafana dashboard accuracy and signal coverage gapshighneeds-triageTSO-0023Separate durable evidence state from poll cursorshighconfigurationobservabilityTSO-0024Support per-tag metrics port overrides in node-metrics discoverymediumneeds-triageTSO-0025Migrate the repo task surface to just and retire Makefiles and ad-hoc scriptsmediumwave:2-fleetTSO-0026Add config.schema.json to the advertised regenerate-everything pathlowneeds-triageTSO-0027Make each generated artifact its own just recipe and retire the script's target dispatchlowneeds-triageTSO-0088Clarify App.Close flow-store ownership outside Run shutdownlowneeds-triageTSO-0110Shard and compress Kubernetes checkpoint storage to remove the 1 MiB ceilinghighTSO-0111Reject a Kubernetes checkpoint configuration that cannot fit its ConfigMaphighTSO-0112Eliminate the load-dependent CI test flakes that make the gate unreliablehighTSO-0113Default the coordination namespace to the release namespace, not defaultmediumTSO-0114Reconcile rollback-era legacy checkpoint writes on re-upgrademediumTSO-0115Prove coordinated Kubernetes mode on the lab, then revertmediumTSO-0116Adjudicate the clientlib-drift alert that fires while both matrix legs passlowTSO-0117Convert the inventoried timeout-only deadlock guards to deterministic barriersmediumTSO-0118Reconcile a legacy checkpoint object whose migration marker an older release removedhighTSO-0119Standby replicas record coordination telemetry that the Prometheus pull path never serveshighneeds-triageTSO-0120Ship alert rules for the Lease coordination feature, which has nonemediumneeds-triageTSO-0121Guard the unsuffixed OpenTelemetry instrumentation scope against a future majormediumneeds-triageTSO-0122Keep alert policy counts and recommended-profile prose truthfullowneeds-triageTSO-0123Allow deployment verification to select a gcx context explicitlyUnspecifiedneeds-triageTSO-0124Alert when an enabled receiver is fail-closed by missing credentialsmediumneeds-triageTSO-0125Alert on Prometheus gather errors that silently omit seriesmediumneeds-triageTSO-0126Alert on sustained processor queue dropsmediumneeds-triageTSO-0127Alert on consecutive profiling upload failureslowneeds-triageTSO-0128Alert on Kubernetes audit schema driftmediumneeds-triageTSO-0129Reject coordination timings that client-go cannot runhighneeds-triageTSO-0130Self-fence promptly when a live coordination Lease is deleted or replacedhighneeds-triageTSO-0131Emit a coordination handover counter so leadership flapping can be alerted truthfullymediumneeds-triageTSO-0132Prove the coordinated Kubernetes surface live: standby scrapes, demotion, Lease loss, clock skewhighneeds-triageTSO-0133Review the coordination alert firing history and decide paging, due 2026-09-11lowTSO-0134PAM config changes never reach the audit-changes metricmediumTSO-0135Assert the pinned golangci-lint and govulncheck versions at run time, the way the helm generators domediumTSO-0136Attribute PAM telemetry to an operator-named tailnet via pam.tailnetmediumTSO-0137Opt-in per-session PAM log record, off by default and pii_filter-gatedlowTSO-0138Discovery sweep: inventory uncollected API surfaces and dead shipped signals to refill the boardmediumTSO-0139Expose safe log-stream destination configuration alongside delivery healthmediumTSO-0140Give the lab deployment a rollout trigger so it tracks releases instead of a stale main digestmediumTSO-0141Bring the lab configuration up to the shipped feature set and adjudicate the zero-sample signal familiesmediumTSO-0142Emit the modeled capability scope preflight metric at runtimemediumTSO-0143Give coordinated replicas per-replica persistent state so the lab can run two replicasmediumTSO-0144Route receiver and admin traffic by a leader pod label so standbys can be ReadyhighTSO-0145Audit v5 documentation against shipped features and remove AI prosehighTSO-0146Promotion-time ingress WAL replay blacks out the leaderhighTSO-0147Resolve recurring readiness failures and ingress WAL drain backloghighTSO-0148Decouple ingress WAL completion from metric collection to preserve configured DPMhighTSO-0149Build local WAL throughput and configuration comparison harnessmediumTSO-0150Diversify Grafana dashboard visualisationshigh