Integrating ClickHouse with Google Meet
The Google Meet registry item copies raw resource readers with a 24-hour lookback and independently journaled discovery-page checkpoints into a chkit project.
Install
Section titled “Install”bunx chkit add google-meet --with-testsbunx chkit checkbunx chkit generate --name add_google_meetbunx chkit migrate --applybunx chkit ingest run --tag provider:google-meetbunx chkit ingest status --tag provider:google-meetSet GOOGLE_MEET_ACCESS_TOKEN with meetings.space.readonly access. Review identity, lookback, schema placement and budgets in src/integrations/google-meet/config.ts. Keep identity tied to one Google account; these list responses do not identify the authenticated account. Coverage depends on that user’s access. The reader does not refresh tokens or export an entire workspace automatically. Schema imports do not call Google. See authorization.
| Resource | Default ClickHouse table | Records synced | API reference |
|---|---|---|---|
Conferences (conferences) | google_meet_conferences_raw | Conference records visible to the access token | GET/conferenceRecords |
Transcripts (transcripts) | google_meet_transcripts_raw | Native transcript metadata for independently discovered recent and ongoing conferences | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/transcripts |
Transcript entries (transcript-entries) | google_meet_transcript_entries_raw | Individually keyed transcript entries for independently discovered recent and ongoing conferences | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/transcripts, GET/conferenceRecords/{conferenceRecord}/transcripts/{transcript}/entries |
Participants (participants) | google_meet_participants_raw | Native attendance records, including signed-in, anonymous, and phone participants | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/participants |
Participant sessions (participant-sessions) | google_meet_participant_sessions_raw | Individually keyed join/leave sessions with conference and participant names | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/participants, GET/conferenceRecords/{conferenceRecord}/participants/{participant}/participantSessions |
Recordings (recordings) | google_meet_recordings_raw | Recording metadata and native Drive destination references; no file downloads | GET/conferenceRecords, GET/conferenceRecords/{conferenceRecord}/recordings |
Pipeline and streams
Section titled “Pipeline and streams”One installation pipeline contains six independent streams: conferences, participants, participant-sessions, recordings, transcripts, and transcript-entries. Each owns its raw destination and journal progress. Every child stream discovers its own conference IDs; entries discover transcripts and sessions discover participants. Resources run independently in any order.
bunx chkit ingest run --tag provider:google-meet --tag resource:transcript-entriesbunx chkit ingest run --tag provider:google-meet --tag resource:participant-sessionssourceId scopes raw identities and queries; streamPrefix determines journal IDs. Another installation needs separate values. The factory snapshots supplied configuration and injectable request dependencies. Edit database before schema imports to change setup-time table placement. The default pipeline runs one stream and request at a time.
Sync behavior
Section titled “Sync behavior”Initial discovery covers the preceding 24 hours by conference end time. Later windows start at the completed watermark minus lookbackHours (24 by default), rereading recent calls and catching up after missed runs. Ongoing calls are discovered separately using end_time IS NULL, including long-running calls started before the lookback.
Children of selected conferences are read completely with ordinary pagination. No conference catalog or child queue survives a completed window. Changes that appear after their conference leaves that window require a wider configured lookback or explicit backfill. There is no automatic historical reconciliation. Conference end times are not child update times; this is bounded recent polling. See conference filters and artifact behavior.
Checkpoints and recovery
Section titled “Checkpoints and recovery”The saved state contains scope, watermark, optional historical bounds, and the unfinished discovery window/phase/page token. Bounds are saved before the first request. A discovery page’s children finish before a state-only marker advances the parent-page cursor. The executor commits it after all preceding rows land.
An interrupted page repeats its child reads from the beginning under the same bounds; completed discovery pages remain skipped. No child cursor, cached page, or child completion index is persisted. Configure small discovery pages and enough budget to finish one page’s children plus the checkpoint marker. Defaults are pageSize: 1 and maxChunks: 200; increase the cap or --max-duration for unusually large conferences. Native JSON requires ClickHouse 25.3 or later.
Rejected tokens restart their collection once. Discovery recovery debt persists across resumed executions until the phase finishes. Child collections restart naturally when their parent page replays. Repeated rejection and cyclic/malformed pagination fail visibly. Strategy version 2 rejects incompatible saved state; scope changes require deliberate migration or a new streamPrefix. Schedule externally with one ingestion process per target.
Raw updates and historical ranges
Section titled “Raw updates and historical ranges”Raw identity is [sourceId, provider.name]. Children add conference_name; sessions add participant_name, and entries add transcript_name. Join native resource names and derive attendance/transcript views later in ClickHouse.
Participants and sessions preserve attendance identities and join/leave records. Recordings preserve metadata and Drive references without downloading media. Transcripts preserve metadata/document references, and entries preserve speech segments. API entries do not capture later edits to the separate Google Docs transcript.
Each observation inserts under its stable identity into ReplacingMergeTree(_chkit_ingested_at). A later transcript state replaces an earlier state; an empty list creates no placeholder that could block a generated transcript. A body hash or insert-only staging is unnecessary. Query with FINAL for latest observations. Missing or expired records are not inferred deletions; archived raw rows remain. See conference retention.
Backfills select conference end times and skip ongoing calls. Child lists are complete for selected conferences. Use the same ID and bounds to resume; changed bounds require a new ID, and expired API history is unavailable.
bunx chkit ingest run --tag provider:google-meet --backfill october --from 2026-10-01T00:00:00Z --to 2026-10-05T00:00:00ZVersion 0.2.0 adds participant, session, recording and entry destinations while preserving conference/transcript tables. Entries become separate rows. Source-scoped row IDs do not replace legacy provider-name IDs; archive or migrate legacy rows deliberately. The installed README explains this upgrade.
Fixture verification
Section titled “Fixture verification”bun test src/integrations/google-meet/tests/basic.test.tsFixtures cover independent resources, coarse replay/resume, lost sink acknowledgement, changed transcript content under stable IDs, recent late artifacts, empty pages, bounded recovery, fixed windows, scope and backfill isolation.
Changelog
Section titled “Changelog”Version 0.2.0
- Sync six independent raw resources in one installation pipeline: conferences, participants, participant sessions, recordings, transcript metadata, and individual transcript entries.
- Default initial discovery and subsequent overlap to 24 hours, including ongoing calls; reread child collections without retaining conference catalogs or child work queues.
- Checkpoint completed discovery pages after their child rows land. Replay unfinished pages with ordinary pagination, fixed bounds, source scope, and bounded token recovery.
- Keep recording metadata and Drive references without downloading files; keep child rows separate with provider-name join keys and source-scoped row IDs. Migrate legacy transcript identities and embedded entries deliberately when upgrading.
- Separate editable configuration, injectable requests, and resource readers. Broader artifact coverage uses an explicit lookback or isolated backfill rather than automatic historical reconciliation.
Version 0.1.0
- Introduce raw conference and transcript ingestion with transcript entries nested in each transcript observation.