Completed work
tailscale2otel
152 completed tasks.
TSO-0150Diversify Grafana dashboard visualisationsTSO-0149Build local WAL throughput and configuration comparison harnessTSO-0147Resolve recurring readiness failures and ingress WAL drain backlogTSO-0146Promotion-time ingress WAL replay blacks out the leaderTSO-0145Audit v5 documentation against shipped features and remove AI proseTSO-0144Route receiver and admin traffic by a leader pod label so standbys can be ReadyTSO-0033Design and ship an HA / multi-replica deployment storyTSO-0142Emit the modeled capability scope preflight metric at runtimeTSO-0143Give coordinated replicas per-replica persistent state so the lab can run two replicasTSO-0140Give the lab deployment a rollout trigger so it tracks releases instead of a stale main digestTSO-0139Expose safe log-stream destination configuration alongside delivery healthTSO-0141Bring the lab configuration up to the shipped feature set and adjudicate the zero-sample signal familiesTSO-0135Assert the pinned golangci-lint and govulncheck versions at run time, the way the helm generators doTSO-0136Attribute PAM telemetry to an operator-named tailnet via pam.tailnetTSO-0137Opt-in per-session PAM log record, off by default and pii_filter-gatedTSO-0138Discovery sweep: inventory uncollected API surfaces and dead shipped signals to refill the boardTSO-0134PAM config changes never reach the audit-changes metricTSO-0036PAM telemetry collector (Border0 API)TSO-0130Self-fence promptly when a live coordination Lease is deleted or replacedTSO-0131Emit a coordination handover counter so leadership flapping can be alerted truthfullyTSO-0132Prove the coordinated Kubernetes surface live: standby scrapes, demotion, Lease loss, clock skewTSO-0124Alert when an enabled receiver is fail-closed by missing credentialsTSO-0125Alert on Prometheus gather errors that silently omit seriesTSO-0126Alert on sustained processor queue dropsTSO-0127Alert on consecutive profiling upload failuresTSO-0128Alert on Kubernetes audit schema driftTSO-0129Reject coordination timings that client-go cannot runTSO-0119Standby replicas record coordination telemetry that the Prometheus pull path never servesTSO-0120Ship alert rules for the Lease coordination feature, which has noneTSO-0121Guard the unsuffixed OpenTelemetry instrumentation scope against a future majorTSO-0122Keep alert policy counts and recommended-profile prose truthfulTSO-0123Allow deployment verification to select a gcx context explicitlyTSO-0094Fail startup for network receivers missing credentials in the next majorTSO-0118Reconcile a legacy checkpoint object whose migration marker an older release removedTSO-0115Prove coordinated Kubernetes mode on the lab, then revertTSO-0113Default the coordination namespace to the release namespace, not defaultTSO-0114Reconcile rollback-era legacy checkpoint writes on re-upgradeTSO-0116Adjudicate the clientlib-drift alert that fires while both matrix legs passTSO-0117Convert the inventoried timeout-only deadlock guards to deterministic barriersTSO-0112Eliminate the load-dependent CI test flakes that make the gate unreliableTSO-0110Shard and compress Kubernetes checkpoint storage to remove the 1 MiB ceilingTSO-0111Reject a Kubernetes checkpoint configuration that cannot fit its ConfigMapTSO-0093Investigate high stream flow ingest-event-age p95TSO-0104Persistent flow store fails to open after RC rolloutTSO-0105Grafana alert rules remain unevaluated after publicationTSO-0106Stream flow capture-delay telemetry is absentTSO-0107Lease-based active-passive coordination (HA phase A1)TSO-0108Kubernetes checkpoint store backend (HA phase A2)TSO-0109Helm chart support for coordinated multi-replica deploymentTSO-0095Validate Wave 3 live on the lab deploymentTSO-0096Shard the CodeRabbit pre-commit review so a wave-sized diff can be reviewedTSO-0097Decide precedence between env-supplied secrets and their _file siblingsTSO-0098Webhook router caches its tokenless and auth-mix decision across secret rotationTSO-0099Services collector reports a complete host snapshot after cancelled dispatchTSO-0100Keep the docker-compose deployment path validated by CI now camden is retiredTSO-0101Reject an already-cancelled request before it consumes an admission slotTSO-0103Restyle the embedded console onto design system v2 (family standard-setter)TSO-0102Split the shared TS_WIF secret: one name, two incompatible scope requirementsTSO-0034Org auto-discovery of tailnets via the alpha Organizations APITSO-0039Posture attribute values: compliance gauges with cardinality capsTSO-0040Workload Identity Federation as an exporter auth methodTSO-0052Bound the N+1 per-device subrequests in the devices/services collectorsTSO-0055Classify nodemetrics scrape failures beyond transient_failureTSO-0056Scale the scheduler stagger window with deployment sizeTSO-0057Per-type intervals for logstream status probesTSO-0058Hot-reload receiver secrets (streaming token, webhook secret)TSO-0059Per-tailnet admission fairness on shared receiversTSO-0061First-class ingest-lag signal per source and signal typeTSO-0063Act on device-cache staleness during control-plane outagesTSO-0067Cost forecast for expensive collector knobsTSO-0068Delta-temporality escape hatch for OTLP metrics exportTSO-0069Promote the OTLP-outage diagnostics summary interval to configTSO-0070Warn on gRPC exporter with CA rotation but no reconnection periodTSO-0071Per-tailnet Tailscale API rate-limit utilization gaugeTSO-0072Rate-limit admin authentication failuresTSO-0073Mutual TLS on the admin listenerTSO-0074Include a bounded recent-log tail in the support bundleTSO-0075Consolidated durable-storage health view on the status pageTSO-0076Debounce checkpoint writes across collectorsTSO-0079Env-var injection for list-valued credentials (tailnets, routes, targets)TSO-0080Surface zero-traffic receiver misconfiguration without breaking startupTSO-0081Version-check fail-open visibility and Headscale defaultTSO-0082Flow store disk reclamation and journal observabilityTSO-0083Headscale and multi-tailnet starter configs in examples/configTSO-0084Extend the compose secrets template to objectstore and Headscale credentialsTSO-0085Complete the NetworkPolicy egress guidance (MaxMind, FQDN example, sidecar note)TSO-0088Clarify App.Close flow-store ownership outside Run shutdownTSO-0090just verify-deploy reports unreachable when any unrelated gcx context is brokenTSO-0091Dedup youngest-eviction-age gauge latches for the process lifetime with no reset pathTSO-0092Retire the 35-panel ceiling and re-group the whole dashboard tab structureTSO-0038Peer-relay connection dimension on node connectivity metricsTSO-0041Verify flow-log native actor identity fields are decoded, not droppedTSO-0035Graceful handling of new audit-log event families (PAM_*, BORDER0_API, APERTURE_*)TSO-0037Services collector: displayName, tags rollup and NodeID joinTSO-0042Adjudicate the 14 unhandled API response fields in the contract ledgerTSO-0043Grants-aware policy parsing and acls-vs-grants adoption metricTSO-0044Policy file snapshots to Loki (full ACL/grants body as log records)TSO-0045Policy diff log records on revision changeTSO-0046Config-state snapshot family: DNS, settings, webhooks, posture integrations to LokiTSO-0047Device inventory change-log to LokiTSO-0048Grafana annotations from audit events on the generated dashboardsTSO-0049ACL risk findings as structured log recordsTSO-0050Key and user-invite lifecycle timeline eventsTSO-0060WAL fill-percentage gauge and shipped alertTSO-0062TTL reset for schema-drift warning dedup in long-lived processesTSO-0064rdns warm-start snapshot across restartsTSO-0065Youngest-eviction-age gauge on dedup setsTSO-0086Doc and dashboard polish batchTSO-0089Retire or wire the orphaned deploy/grafana/gen/tabs/events.py moduleTSO-0028Fix stale hand-maintained alert-count prose in READMEsTSO-0029Fix stale single-tailnet receiver claim in config.example.yamlTSO-0030WAL replay can double-emit metrics/logs after a crash between apply and commitTSO-0031Headscale custom ip-prefix deployments misclassify tailnet addresses as external and geoip-enrich themTSO-0032Per-runtime shutdown is sequential, contradicting the deliberate parallel-shutdown fixTSO-0051Quiet key-expiry warnings: on-change + daily mode as defaultTSO-0053Cardinality backstop for the posture attribute-namespace wildcardTSO-0054Configurable dedup and seen-set capacities (flow, audit, objectstore)TSO-0066Per-tailnet cardinality limit overridesTSO-0077Headscale HTTP retry/rate-limit config: implement or rejectTSO-0078Document restart-required vs hot-reloadable config keysTSO-0087Classify the three audit enum values the 2026-08-30 spec re-vendor addedTSO-0027Make each generated artifact its own just recipe and retire the script's target dispatchTSO-0026Add config.schema.json to the advertised regenerate-everything pathTSO-0025Migrate the repo task surface to just and retire Makefiles and ad-hoc scriptsTSO-0024Support per-tag metrics port overrides in node-metrics discoveryTSO-0022Fix live Grafana dashboard accuracy and signal coverage gapsTSO-0023Separate durable evidence state from poll cursorsTSO-0021Update check fails permanently: the 64 KiB body cap truncates the GitHub latest-release responseTSO-0020Rotate credentials exposed in an agent transcriptTSO-0008Add an explicit Prometheus-only delivery modeTSO-0007Make the default Prometheus listener safe and scrapeableTSO-0009Add a bounded Prometheus first-scrape checkTSO-0010Make first-run onboarding backend-neutralTSO-0011Ship minimal delivery-specific starter configurationsTSO-0012Revalidate and align the Alloy gateway recipeTSO-0013Document dynamic node-metrics discovery as a first-use pathTSO-0014Make optional ingestion paths easier to chooseTSO-0015Add an operator-first alert profile chooserTSO-0016Add a release-independent upgrade and rollback checklistTSO-0017Fix documentation social metadata ownership and asset checksTSO-0018Decide the supported alerting path for Prometheus-only deploymentsTSO-0019Generate public capability counts from the catalogTSO-0005Persistent flow store refuses every pre-hardening database with no way to migrate itTSO-0001Upgrade Go toolchain to 1.27TSO-0006Migrate to OpenTelemetry v1.46.0 / log v0.22.0 after the log attribute API removalTSO-0003Cut CI job fan-out per PR to drain the Actions queueTSO-0004prometheus.max_requests_in_flight: 0 now crash-loops every config copied from the old exampleTSO-0002Remediate Codex Security scan b3c6de8eTSO-0002.01Harden ingress and local HTTP boundariesTSO-0002.02Bound retained state and protect local persistenceTSO-0002.03Harden outbound credentials transports and TLSTSO-0002.04Close Helm credential and workload identity gaps