Skip to content

Integrating ClickHouse with Linear

The Linear registry item copies independent raw resource readers and ClickHouse destinations into a chkit project.

Terminal window
bunx chkit add linear
bunx chkit check
bunx chkit generate --name add_linear
bunx chkit migrate --apply
bunx chkit ingest run --tag provider:linear

Set LINEAR_API_KEY in the runtime environment before ingestion. Edit queries and raw destinations in src/integrations/linear/sources/ to select fields and configure tables. config.ts supplies source identity and timestamp selection; client.ts handles authenticated requests, and pipeline.ts groups nine independent streams in one workspace pipeline.

createLinearPipeline(config, deps) snapshots configuration and binds injectable HTTP dependencies. The pipeline and all nine table exports remain in index.ts. Destination schemas stay in source modules rather than runtime configuration.

Keep stream IDs and destinations tied to one workspace. The reader does not resolve the token’s workspace identity, so switching workspaces requires new IDs and destinations. Reconcile older observations after adding requested fields.

ResourceDefault ClickHouse tableRecords syncedAPI reference
Issues (issues)linear_issues_rawArchived and active issue fields, relationship IDs, and fully paginated label names as strings.POST/graphql
Comments (comments)linear_comments_rawRaw comments selected by their own update times, with nullable issue and other parent references.POST/graphql
Projects (projects)linear_projects_rawProject metadata and provider relationship references selected by project update times.POST/graphql
Project updates (project_updates)linear_project_updates_rawAuthored progress and health updates with project and author references.POST/graphql
Cycles (cycles)linear_cycles_rawCycle dates, descriptions, and team references selected by cycle update times.POST/graphql
Users (users)linear_users_rawAccessible user metadata, including disabled users, selected by user update times.POST/graphql
Teams (teams)linear_teams_rawAccessible team metadata selected by team update times.POST/graphql
Issue relations (issue_relations)linear_issue_relations_rawComplete accessible relation records with type and both directed issue references; no reliable root change filter exists.POST/graphql
Issue history (issue_history)linear_issue_history_rawComplete accessible history for independently discovered issues; the parent history connection has no timestamp filter.POST/graphql

Issues, comments, projects, project updates, cycles, users, teams, issue relations, and issue history each have their own raw table and checkpoint. Relationships retain provider-shaped ID references. Join resources and calculate metrics later in ClickHouse. Issues retain fully paginated label names as strings; label definitions, workflow states, and project statuses have no separate streams. Bounded inline state/status metadata remains provider fields.

Comments use their own collection, including nullable references for issue and other parent types. History independently discovers all accessible issues before paging their events. Neither reader depends on the issues stream’s execution or checkpoint. Paginated project memberships and other associations are outside the selected resource scope.

Terminal window
bunx chkit ingest run --tag provider:linear --tag resource:comments
bunx chkit ingest run --tag provider:linear --tag resource:issue_history

Repeated tags use AND matching. The default pipeline executes streams sequentially to limit API usage; ordering does not establish a discovery dependency. Expensive history reads can have their own schedule and execution budget.

Seven root resources use server-side bounded updatedAt windows, Relay cursors, five minutes of overlap, and explicit archive inclusion. Users also include disabled accounts. Comments and project updates use their own update times, independently of their parents. Linear pagination and date filtering document these selection controls.

paginate() owns request retries and continuation cycle detection. Readers validate GraphQL errors and reject partial success. Each timestamp checkpoint advances only after the complete window and destination writes succeed, including an empty window. Failed windows replay from their lower bound. Issues complete all label pages before publishing their label strings.

Issue relations lack a reliable root change filter. History is only exposed through each issue’s history connection, with no timestamp filter. These streams use fullSync(): discover their complete accessible scope on each run and journal success after destination acknowledgement, including empty results. Interrupted full reads restart from the beginning. They ignore date bounds and provide no date-range backfill. See Linear’s API schema.

Deleted records and removed relations remain stored. The API does not promise atomic snapshots, and label renames are not assumed to update issue timestamps. Periodic reconciliation refreshes older timestamp-selected observations:

Terminal window
bunx chkit ingest run --tag provider:linear --tag resource:issues --backfill reconcile-2026-10-05 --from 1970-01-01
bunx chkit ingest status --tag provider:linear --json

Use a new backfill ID for each reconciliation. Timestamp replay reads current observations; the history stream preserves accessible provider events. The raw tables require ClickHouse 25.3 or later. See the installed README for configuration and registry installation.

Version 0.2.0 adds eight raw destinations while retaining issue row IDs and the linear.issues stream ID. Existing nested comments in old issue observations remain until those rows are refreshed or migrated deliberately. Initial reads populate the new resource tables; application queries should join those destinations rather than use retained nested comments.

Terminal window
bunx chkit add linear --with-tests
bun test src/integrations/linear/tests/basic.test.ts

Version 0.2.0

  • Sync issues, comments, projects, project updates, cycles, users, and teams through independent overlapping updated-time windows, including archived and disabled resources.
  • Add independent full-sync streams for issue relations and issue history; discover and paginate history for all accessible issues without relying on parent stream progress or modification times.
  • Store each resource in its own raw table with provider IDs and relationship references; retain complete issue labels as strings and leave joins and metrics to ClickHouse.
  • Use shared pagination, validated GraphQL responses, request-level retries and sink-acknowledged completion checkpoints; separate configuration, injectable requests, readers, and pipeline wiring.
  • Existing issue IDs remain stable; migrate retained nested comment observations deliberately and populate the eight new destinations through their own initial syncs.
  • Consume full pagination pages and use shared continuation-cycle detection while preserving independent resource checkpoints and GraphQL response validation.

Version 0.1.0

  • Introduce raw GraphQL issue ingestion.