Completed work

tailscale2otel

152 completed tasks.

TSO-0150Diversify Grafana dashboard visualisationshighTSO-0149Build local WAL throughput and configuration comparison harnessmediumTSO-0147Resolve recurring readiness failures and ingress WAL drain backloghighTSO-0146Promotion-time ingress WAL replay blacks out the leaderhighTSO-0145Audit v5 documentation against shipped features and remove AI prosehighTSO-0144Route receiver and admin traffic by a leader pod label so standbys can be ReadyhighTSO-0033Design and ship an HA / multi-replica deployment storymediumDecision-gatedTSO-0142Emit the modeled capability scope preflight metric at runtimemediumTSO-0143Give coordinated replicas per-replica persistent state so the lab can run two replicasmediumTSO-0140Give the lab deployment a rollout trigger so it tracks releases instead of a stale main digestmediumTSO-0139Expose safe log-stream destination configuration alongside delivery healthmediumTSO-0141Bring the lab configuration up to the shipped feature set and adjudicate the zero-sample signal familiesmediumTSO-0135Assert the pinned golangci-lint and govulncheck versions at run time, the way the helm generators domediumTSO-0136Attribute PAM telemetry to an operator-named tailnet via pam.tailnetmediumTSO-0137Opt-in per-session PAM log record, off by default and pii_filter-gatedlowTSO-0138Discovery sweep: inventory uncollected API surfaces and dead shipped signals to refill the boardmediumTSO-0134PAM config changes never reach the audit-changes metricmediumTSO-0036PAM telemetry collector (Border0 API)mediumDecision-gatedTSO-0130Self-fence promptly when a live coordination Lease is deleted or replacedhighneeds-triageTSO-0131Emit a coordination handover counter so leadership flapping can be alerted truthfullymediumneeds-triageTSO-0132Prove the coordinated Kubernetes surface live: standby scrapes, demotion, Lease loss, clock skewhighneeds-triageTSO-0124Alert when an enabled receiver is fail-closed by missing credentialsmediumneeds-triageTSO-0125Alert on Prometheus gather errors that silently omit seriesmediumneeds-triageTSO-0126Alert on sustained processor queue dropsmediumneeds-triageTSO-0127Alert on consecutive profiling upload failureslowneeds-triageTSO-0128Alert on Kubernetes audit schema driftmediumneeds-triageTSO-0129Reject coordination timings that client-go cannot runhighneeds-triageTSO-0119Standby replicas record coordination telemetry that the Prometheus pull path never serveshighneeds-triageTSO-0120Ship alert rules for the Lease coordination feature, which has nonemediumneeds-triageTSO-0121Guard the unsuffixed OpenTelemetry instrumentation scope against a future majormediumneeds-triageTSO-0122Keep alert policy counts and recommended-profile prose truthfullowneeds-triageTSO-0123Allow deployment verification to select a gcx context explicitlyUnspecifiedneeds-triageTSO-0094Fail startup for network receivers missing credentials in the next majormediumv5 big bang majorTSO-0118Reconcile a legacy checkpoint object whose migration marker an older release removedhighTSO-0115Prove coordinated Kubernetes mode on the lab, then revertmediumTSO-0113Default the coordination namespace to the release namespace, not defaultmediumTSO-0114Reconcile rollback-era legacy checkpoint writes on re-upgrademediumTSO-0116Adjudicate the clientlib-drift alert that fires while both matrix legs passlowTSO-0117Convert the inventoried timeout-only deadlock guards to deterministic barriersmediumTSO-0112Eliminate the load-dependent CI test flakes that make the gate unreliablehighTSO-0110Shard and compress Kubernetes checkpoint storage to remove the 1 MiB ceilinghighTSO-0111Reject a Kubernetes checkpoint configuration that cannot fit its ConfigMaphighTSO-0093Investigate high stream flow ingest-event-age p95mediumWave 5: HA coordination and validation follow-upsTSO-0104Persistent flow store fails to open after RC rollouthighWave 5: HA coordination and validation follow-upsTSO-0105Grafana alert rules remain unevaluated after publicationhighWave 5: HA coordination and validation follow-upsTSO-0106Stream flow capture-delay telemetry is absentmediumWave 5: HA coordination and validation follow-upsTSO-0107Lease-based active-passive coordination (HA phase A1)highWave 5: HA coordination and validation follow-upsTSO-0108Kubernetes checkpoint store backend (HA phase A2)mediumWave 5: HA coordination and validation follow-upsTSO-0109Helm chart support for coordinated multi-replica deploymentmediumWave 5: HA coordination and validation follow-upsTSO-0095Validate Wave 3 live on the lab deploymenthighWave 4: live validation and review follow-upsneeds-triageTSO-0096Shard the CodeRabbit pre-commit review so a wave-sized diff can be reviewedhighWave 4: live validation and review follow-upsneeds-triageTSO-0097Decide precedence between env-supplied secrets and their _file siblingsmediumWave 4: live validation and review follow-upsneeds-triageTSO-0098Webhook router caches its tokenless and auth-mix decision across secret rotationhighWave 4: live validation and review follow-upsneeds-triageTSO-0099Services collector reports a complete host snapshot after cancelled dispatchmediumWave 4: live validation and review follow-upsneeds-triageTSO-0100Keep the docker-compose deployment path validated by CI now camden is retiredmediumWave 4: live validation and review follow-upsneeds-triageTSO-0101Reject an already-cancelled request before it consumes an admission slotlowWave 4: live validation and review follow-upsneeds-triageTSO-0103Restyle the embedded console onto design system v2 (family standard-setter)highWave 4: live validation and review follow-upsdesign-systemTSO-0102Split the shared TS_WIF secret: one name, two incompatible scope requirementshighWave 4: live validation and review follow-upsneeds-triageTSO-0034Org auto-discovery of tailnets via the alpha Organizations APImediumWave 3: API surface expansionTSO-0039Posture attribute values: compliance gauges with cardinality capsmediumWave 3: API surface expansionTSO-0040Workload Identity Federation as an exporter auth methodmediumWave 3: API surface expansionTSO-0052Bound the N+1 per-device subrequests in the devices/services collectorshighWave 4: collector and ingestion resilienceTSO-0055Classify nodemetrics scrape failures beyond transient_failuremediumWave 4: collector and ingestion resilienceTSO-0056Scale the scheduler stagger window with deployment sizelowWave 6: config ergonomics and performanceTSO-0057Per-type intervals for logstream status probeslowWave 6: config ergonomics and performanceTSO-0058Hot-reload receiver secrets (streaming token, webhook secret)highWave 4: collector and ingestion resilienceTSO-0059Per-tailnet admission fairness on shared receiversmediumWave 4: collector and ingestion resilienceTSO-0061First-class ingest-lag signal per source and signal typemediumWave 4: collector and ingestion resilienceTSO-0063Act on device-cache staleness during control-plane outagesmediumWave 4: collector and ingestion resilienceTSO-0067Cost forecast for expensive collector knobsmediumWave 5: operator surface and cost controlsTSO-0068Delta-temporality escape hatch for OTLP metrics exportmediumWave 5: operator surface and cost controlsTSO-0069Promote the OTLP-outage diagnostics summary interval to configlowWave 5: operator surface and cost controlsTSO-0070Warn on gRPC exporter with CA rotation but no reconnection periodlowWave 5: operator surface and cost controlsTSO-0071Per-tailnet Tailscale API rate-limit utilization gaugemediumWave 5: operator surface and cost controlsTSO-0072Rate-limit admin authentication failureshighWave 5: operator surface and cost controlsTSO-0073Mutual TLS on the admin listenermediumWave 5: operator surface and cost controlsTSO-0074Include a bounded recent-log tail in the support bundlemediumWave 5: operator surface and cost controlsTSO-0075Consolidated durable-storage health view on the status pagemediumWave 5: operator surface and cost controlsTSO-0076Debounce checkpoint writes across collectorsmediumWave 6: config ergonomics and performanceTSO-0079Env-var injection for list-valued credentials (tailnets, routes, targets)mediumWave 6: config ergonomics and performanceTSO-0080Surface zero-traffic receiver misconfiguration without breaking startupmediumWave 6: config ergonomics and performanceTSO-0081Version-check fail-open visibility and Headscale defaultlowWave 5: operator surface and cost controlsTSO-0082Flow store disk reclamation and journal observabilitymediumWave 6: config ergonomics and performanceTSO-0083Headscale and multi-tailnet starter configs in examples/configmediumWave 7: deploy and docs polishTSO-0084Extend the compose secrets template to objectstore and Headscale credentialslowWave 7: deploy and docs polishTSO-0085Complete the NetworkPolicy egress guidance (MaxMind, FQDN example, sidecar note)lowWave 7: deploy and docs polishTSO-0088Clarify App.Close flow-store ownership outside Run shutdownlowneeds-triageTSO-0090just verify-deploy reports unreachable when any unrelated gcx context is brokenmediumWave 5: operator surface and cost controlsneeds-triageTSO-0091Dedup youngest-eviction-age gauge latches for the process lifetime with no reset pathmediumWave 4: collector and ingestion resilienceneeds-triageTSO-0092Retire the 35-panel ceiling and re-group the whole dashboard tab structurehighWave 7: deploy and docs polishneeds-triageTSO-0038Peer-relay connection dimension on node connectivity metricsmediumWave 3: API surface expansionTSO-0041Verify flow-log native actor identity fields are decoded, not droppedmediumWave 3: API surface expansionTSO-0035Graceful handling of new audit-log event families (PAM_*, BORDER0_API, APERTURE_*)mediumWave 3: API surface expansionTSO-0037Services collector: displayName, tags rollup and NodeID joinmediumWave 3: API surface expansionTSO-0042Adjudicate the 14 unhandled API response fields in the contract ledgerlowWave 3: API surface expansionTSO-0043Grants-aware policy parsing and acls-vs-grants adoption metricmediumWave 3: API surface expansionTSO-0044Policy file snapshots to Loki (full ACL/grants body as log records)mediumWave 2: Loki state-snapshot familyTSO-0045Policy diff log records on revision changemediumWave 2: Loki state-snapshot familyTSO-0046Config-state snapshot family: DNS, settings, webhooks, posture integrations to LokimediumWave 2: Loki state-snapshot familyTSO-0047Device inventory change-log to LokimediumWave 2: Loki state-snapshot familyTSO-0048Grafana annotations from audit events on the generated dashboardslowWave 2: Loki state-snapshot familyTSO-0049ACL risk findings as structured log recordslowWave 2: Loki state-snapshot familyTSO-0050Key and user-invite lifecycle timeline eventslowWave 2: Loki state-snapshot familyTSO-0060WAL fill-percentage gauge and shipped alertmediumWave 4: collector and ingestion resilienceTSO-0062TTL reset for schema-drift warning dedup in long-lived processeslowWave 4: collector and ingestion resilienceTSO-0064rdns warm-start snapshot across restartsmediumWave 4: collector and ingestion resilienceTSO-0065Youngest-eviction-age gauge on dedup setslowWave 4: collector and ingestion resilienceTSO-0086Doc and dashboard polish batchlowWave 7: deploy and docs polishTSO-0089Retire or wire the orphaned deploy/grafana/gen/tabs/events.py modulemediumWave 2: Loki state-snapshot familyneeds-triageTSO-0028Fix stale hand-maintained alert-count prose in READMEsmediumWave 1: bugs and guardrail foundationsTSO-0029Fix stale single-tailnet receiver claim in config.example.yamllowWave 1: bugs and guardrail foundationsTSO-0030WAL replay can double-emit metrics/logs after a crash between apply and commitmediumWave 1: bugs and guardrail foundationsTSO-0031Headscale custom ip-prefix deployments misclassify tailnet addresses as external and geoip-enrich themmediumWave 1: bugs and guardrail foundationsTSO-0032Per-runtime shutdown is sequential, contradicting the deliberate parallel-shutdown fixlowWave 1: bugs and guardrail foundationsTSO-0051Quiet key-expiry warnings: on-change + daily mode as defaulthighWave 1: bugs and guardrail foundationsTSO-0053Cardinality backstop for the posture attribute-namespace wildcardmediumWave 1: bugs and guardrail foundationsTSO-0054Configurable dedup and seen-set capacities (flow, audit, objectstore)mediumWave 1: bugs and guardrail foundationsTSO-0066Per-tailnet cardinality limit overrideshighWave 1: bugs and guardrail foundationsTSO-0077Headscale HTTP retry/rate-limit config: implement or rejecthighWave 1: bugs and guardrail foundationsTSO-0078Document restart-required vs hot-reloadable config keysmediumWave 1: bugs and guardrail foundationsTSO-0087Classify the three audit enum values the 2026-08-30 spec re-vendor addedhighWave 1: bugs and guardrail foundationsneeds-triageTSO-0027Make each generated artifact its own just recipe and retire the script's target dispatchlowneeds-triageTSO-0026Add config.schema.json to the advertised regenerate-everything pathlowneeds-triageTSO-0025Migrate the repo task surface to just and retire Makefiles and ad-hoc scriptsmediumwave:2-fleetTSO-0024Support per-tag metrics port overrides in node-metrics discoverymediumneeds-triageTSO-0022Fix live Grafana dashboard accuracy and signal coverage gapshighneeds-triageTSO-0023Separate durable evidence state from poll cursorshighconfigurationobservabilityTSO-0021Update check fails permanently: the 64 KiB body cap truncates the GitHub latest-release responsehighTSO-0020Rotate credentials exposed in an agent transcripthighneeds-triagesecurityTSO-0008Add an explicit Prometheus-only delivery modehighuser-friendlinessneeds-triageuser-friendlinessTSO-0007Make the default Prometheus listener safe and scrapeablehighuser-friendlinessneeds-triageuser-friendlinessTSO-0009Add a bounded Prometheus first-scrape checkmediumuser-friendlinessneeds-triageuser-friendlinessTSO-0010Make first-run onboarding backend-neutralhighuser-friendlinessneeds-triageuser-friendlinessTSO-0011Ship minimal delivery-specific starter configurationshighuser-friendlinessneeds-triageuser-friendlinessTSO-0012Revalidate and align the Alloy gateway recipemediumuser-friendlinessneeds-triageuser-friendlinessTSO-0013Document dynamic node-metrics discovery as a first-use pathmediumuser-friendlinessneeds-triageuser-friendlinessTSO-0014Make optional ingestion paths easier to choosemediumuser-friendlinessneeds-triageuser-friendlinessTSO-0015Add an operator-first alert profile choosermediumuser-friendlinessneeds-triageuser-friendlinessTSO-0016Add a release-independent upgrade and rollback checklistmediumuser-friendlinessneeds-triageuser-friendlinessTSO-0017Fix documentation social metadata ownership and asset checkslowuser-friendlinessneeds-triageuser-friendlinessTSO-0018Decide the supported alerting path for Prometheus-only deploymentsmediumuser-friendlinessneeds-triageuser-friendlinessTSO-0019Generate public capability counts from the cataloglowuser-friendlinessneeds-triageuser-friendlinessTSO-0005Persistent flow store refuses every pre-hardening database with no way to migrate itUnspecifiedTSO-0001Upgrade Go toolchain to 1.27UnspecifiedTSO-0006Migrate to OpenTelemetry v1.46.0 / log v0.22.0 after the log attribute API removalUnspecifiedTSO-0003Cut CI job fan-out per PR to drain the Actions queueUnspecifiedTSO-0004prometheus.max_requests_in_flight: 0 now crash-loops every config copied from the old exampleUnspecifiedTSO-0002Remediate Codex Security scan b3c6de8ehighTSO-0002.01Harden ingress and local HTTP boundarieshighTSO-0002.02Bound retained state and protect local persistencehighTSO-0002.03Harden outbound credentials transports and TLShighTSO-0002.04Close Helm credential and workload identity gapshigh