Skip to content

Fix telemetry endpoint, deploy collector on telemetry.pgflow.dev, and release patch #691

Description

@jumski

Summary

Deploy the pgflow telemetry collector behind https://telemetry.pgflow.dev, correct the hardcoded database and CLI sender URLs, and ship a patch release after the endpoint works. Version 0.17.1 points both senders at https://pgflow-telemetry.workers.dev, which does not resolve. Deploying the existing Worker alone cannot repair that URL.

User report or idea

After releasing 0.17.1, the user requested deployment of the Cloudflare collector. Deployment preflight exposed the endpoint mismatch. The user first asked whether this could be fixed without another release, then requested a plan for a new worktree, a corrective release, and pgflow domain/Cloudflare setup together.

This issue is the handoff for that new worktree. Do not create the worktree, deploy infrastructure, or publish a release as part of issue creation. Implement in a fresh worktree from current origin/main, not directly on the bot-owned changeset-release/main branch.

For release announcements, lead with the fix and keep telemetry detail small; link to the telemetry reference or news page rather than repeating the full explanation. The existing 0.17.1 Discord/X notes already follow that preference.

Evidence supplied

Hardcoded senders

pkgs/cli/src/commands/install/report-install-telemetry.ts:

const ENDPOINT = 'https://pgflow-telemetry.workers.dev/';

pkgs/core/schemas/0135_function_report.sql:

    v_request_id := net.http_post(
      url => 'https://pgflow-telemetry.workers.dev',
      body => v_payload,
      headers => jsonb_build_object('Content-Type', 'application/json'),
      timeout_milliseconds => 5000
    );

The same URL appears in the already-released migration pkgs/core/supabase/migrations/20260920093533_pgflow_telemetry.sql.

Existing collector configuration

apps/telemetry-worker/wrangler.toml:

name = "pgflow-telemetry"
main = "src/index.ts"
compatibility_date = "2026-09-01"

[[analytics_engine_datasets]]
binding = "PGFLOW_TELEMETRY"
dataset = "pgflow_telemetry"

# No request logging: bodies and transport metadata are never stored.
[observability]
enabled = false

Deployment preflight on 2026-10-01:

  • wrangler whoami authenticated using the existing environment token. Account identifiers and user details are omitted from this public issue; no credential values are included.
  • The read-only Cloudflare account subdomain lookup returned jumski, making the default Worker URL https://pgflow-telemetry.jumski.workers.dev, not the released URL.
  • The Worker lookup returned HTTP 404:
{
  "result": null,
  "success": false,
  "errors": [
    {
      "code": 10007,
      "message": "This Worker does not exist on your account."
    }
  ],
  "messages": []
}
  • DNS lookup for pgflow-telemetry.workers.dev and telemetry.pgflow.dev failed with:
[Errno -2] Name or service not known
  • pnpm --dir apps/telemetry-worker exec wrangler deploy --dry-run succeeded, bundling the Worker and listing the PGFLOW_TELEMETRY / pgflow_telemetry Analytics Engine binding. No live deployment occurred.
  • pnpm nx run-many -t test typecheck -p @pgflow/telemetry-worker passed: 27 tests and TypeScript checking.

Published baseline

Version Packages PR #687 merged as 1f8edde03d7cc6bc7985f6f74af70adf04024dca. The Release workflow succeeded. Registry checks confirmed version 0.17.1 for npm packages @pgflow/core, @pgflow/dsl, @pgflow/client, and pgflow, plus JSR @pgflow/edge-worker. The first client lookup failed; a later lookup found 0.17.1 and its latest tag. Registry propagation was the suspected cause, not independently proved.

Investigation and findings

  • Cloudflare documents Worker URLs as <YOUR_WORKER_NAME>.<YOUR_SUBDOMAIN>.workers.dev: https://developers.cloudflare.com/workers/configuration/routing/workers-dev/ . The published URL omits the account subdomain. We do not control the workers.dev DNS zone and cannot add a redirect at the broken hostname through our pgflow zone.
  • Both senders hardcode the endpoint; neither offers a deployment-time endpoint override. Changing collector hosting or documentation cannot repair already-installed senders.
  • CLI telemetry catches send failures and uses a 500 ms timeout; installation does not fail because the collector is unreachable.
  • The database queues an asynchronous pg_net request with a 5-second timeout, records an audit row, and does not inspect delivery status. An audit row or sent: return value is not proof of successful ingestion. A recorded day is skipped on subsequent calls.
  • No Worker was deployed during this investigation. Authentication and dry-run success do not establish production deploy permissions, zone ownership, certificate readiness, or Analytics Engine availability.
  • The current collector validates a closed schema, limits bodies to 2048 bytes and contributions to 64, rejects invalid/duplicate entries before writes, and writes one Analytics Engine point per contribution. Preserve those safeguards and disabled request logging.

Proposed solution or design

1. Start in a fresh worktree and establish the checks

  • Start from current origin/main (0.17.1 merged); suggested branch: fix/telemetry-collector-endpoint.
  • Follow repository worktree lifecycle rules and read affected package instructions. Load the pgflow development, schema, pgTAP, migration, publishing, and documentation skills where applicable.
  • Inspect Nx targets and record the focused-to-final check ladder. Use the repository-managed fresh test environment before implementation; do not reset unrelated Supabase projects.

2. Set up the stable pgflow endpoint and deploy the collector

  • Use https://telemetry.pgflow.dev/ for both senders, avoiding an account-specific public API address.
  • Check the existing pgflow.dev Cloudflare zone, DNS records, Worker deployments, and token permissions before writes. Do not replace website DNS, transfer the domain, rename the account subdomain, or change nameservers as an incidental step.
  • Configure a Worker Custom Domain for telemetry.pgflow.dev using the existing apps/telemetry-worker project and dataset binding. Check Analytics Engine access and TLS readiness. Keep request observability disabled.
  • Run the collector tests, typecheck, and deployment dry-run. Deploy using the repository's existing tooling after confirming the target resources.
  • Check rejection behavior without writing data (GET 405; malformed JSON 400). For ingestion proof, explicitly approve and document one synthetic contribution before writing it to production Analytics Engine; check the stored point if query permissions allow. Do not send a real user's payload or call an audit row proof of delivery.
  • Make the endpoint live before publishing corrected senders.

3. Fix both senders and generate an additive upgrade migration

  • Change the CLI constant and canonical SQL schema to the custom-domain URL. Match slash conventions in assertions.
  • Generate a new migration replacing pgflow_telemetry.report() for existing installations. Do not rewrite the released 20260920093533_pgflow_telemetry.sql or its history just to remove the old hostname.
  • Update regression tests and current endpoint documentation. Inventory all tracked references; retain the historical released migration as intentional evidence of the previous endpoint.
  • Relevant files:
    • pkgs/cli/src/commands/install/report-install-telemetry.ts
    • pkgs/cli/__tests__/commands/install/report-install-telemetry.test.ts
    • pkgs/core/schemas/0135_function_report.sql
    • pkgs/core/supabase/tests/telemetry/report.test.sql
    • apps/telemetry-worker/src/index.test.ts
    • pkgs/website/src/content/docs/reference/telemetry.mdx
  • Preserve opt-out, local/CI suppression, timeout limits, non-fatal failures, payload validation, and privacy behavior. The migration must not silently re-enable telemetry.
  • Verify both a fresh install and an upgrade from 0.17.1 use the corrected SQL function. Follow the schema-first/generated-migration workflow, including migration and generated-type verification.

4. Ship and verify the patch release

  • Add patch changesets for the affected publishable packages (@pgflow/core and pgflow). Preserve the existing fixed version group; the private collector is not a public package.
  • Target 0.17.2 if it remains the next patch when implementation completes; let Changesets resolve the actual version.
  • Open the fix PR and merge only after the applicable tests and CI pass. Let the bot regenerate Version Packages; do not manually maintain its branch.
  • Merge Version Packages only when green and release authorization remains current. Watch main's build/deployment checks and Release workflow.
  • Verify the actual versions and latest tags directly on npm and JSR, then check the production collector endpoint again. Report any missing permission or ingestion evidence honestly.
  • Document that existing installations need the new migration and CLI update; merely deploying the endpoint does not repair 0.17.1.

Alternatives discussed

  • Deploy only: possible, but does not change the hostname used by 0.17.1.
  • No immediate package release: deploy the collector, provide an optional manual SQL replacement, and defer the CLI fix to a later release. This repairs only databases whose owners apply the SQL; it does not repair every installed CLI. The user subsequently requested a corrective release instead.
  • Default workers.dev address: technically suitable for the collector but account-specific. Prefer the stable pgflow custom domain. Renaming the account subdomain does not remove the Worker-name component of a normal workers.dev URL.

Acceptance criteria

  • https://telemetry.pgflow.dev/ resolves with valid TLS and routes to the collector, with the intended Analytics Engine binding and request logging disabled. Production ingestion evidence or its precise remaining permission blocker is recorded.
  • CLI and database senders use the same canonical endpoint; tests and current docs agree. The already-released migration remains unchanged and a new generated migration upgrades 0.17.1 installations.
  • Collector, CLI, SQL, migration, and relevant integration checks pass. Existing opt-out and failure behavior remain intact; the upgrade does not enable telemetry for opted-out databases.
  • A corrective patch release completes through green PR/main/release checks, with npm and JSR versions and tags checked directly.
  • The handoff records infrastructure changes, deployed Worker version, endpoint checks, package versions, and upgrade requirements without exposing credentials or claiming lost telemetry was recovered.

Related work

Open and closed issue searches for telemetry, collector, workers.dev, Cloudflare, and telemetry.pgflow.dev found no matching ordinary issue at planning time. This issue is not attached to a Delivery or roadmap.

Open questions and risks

  • Is the existing pgflow.dev zone active in the authenticated account, and does the token allow Worker Custom Domains, deploys, and Analytics Engine access? These were not checked yet.
  • A Cloudflare Custom Domain/DNS change and a synthetic production analytics write need explicit target/impact confirmation before execution. Use a human setup step only for permissions/actions the agent cannot perform.
  • Existing sent_reports rows may represent failed delivery to the old endpoint and prevent same-day resend. No automatic replay/backfill was requested; do not delete audit rows or replay stored telemetry without a separate decision.
  • Existing 0.17.1 installs will continue using the dead endpoint until owners upgrade or apply an approved manual patch.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions