Skip to content

Integrating ClickHouse with Circleback

The Circleback integration runs five independently selectable raw resource streams within one account pipeline.

Terminal window
bunx chkit add circleback
bunx chkit check
bunx chkit generate --name add_circleback
bunx chkit migrate --apply
bunx chkit ingest run --tag provider:circleback

Set CIRCLEBACK_API_KEY in the runtime environment. Create a key through Circleback Settings → API keys; see the authentication reference. Native JSON requires ClickHouse 25.3 or later.

Edit src/integrations/circleback/config.ts for a stable sourceId, meeting ownership, and setup-time database. ownership: 'All' includes shared meetings accessible to the authenticated user. People, companies, and action items use the provider’s own authenticated-user collections; these endpoints do not promise a full organization export.

pipeline.ts composes the five resources, sources/ owns raw schemas and readers, and client.ts handles requests and Link pagination. The factory createCirclebackPipeline(config, deps) binds reader configuration and optional HTTP dependencies. Runtime factories use existing exported schemas; storage changes belong before schema discovery and migration generation.

ResourceDefault ClickHouse tableRecords syncedAPI reference
Meetings (meetings)circleback_meetings_rawAccessible meeting provider objects and notes, with tag names as string arrays.GET/meetings
Meeting transcripts (meeting_transcripts)circleback_meeting_transcripts_rawWhole transcript snapshots per meeting and separate latest read outcomes; availability failures preserve earlier successful text.GET/meetings, GET/meeting/{meetingId}/transcript
Action items (action_items)circleback_action_items_rawRaw action items for any assignee, fully paginated across PENDING and DONE completion statuses.GET/action-items
People (people)circleback_people_rawComplete same-person listing and detail fields, preserving external references and native person/company IDs.GET/people, GET/person/{profileId}
Companies (companies)circleback_companies_rawDomain-keyed company listing and detail fields with a nullable, corroborated native numeric company ID for joins.GET/companies, GET/company/{domain}

Meetings, meeting transcripts, action items, people, and companies have separate streams, tables, and journal completion. Each runs alone and owns any required discovery. The transcript reader lists its own eligible meetings; selecting transcripts does not depend on a prior meetings run, and transcript failure does not block raw meeting publication.

Raw provider fields remain under data with source_id metadata. Provider-returned meeting attendee/action-item summaries and recording URLs remain intact; readers do not assemble a meeting by fetching its child streams. The requested tag projection stores all returned tag names as string[], without a tag-definition stream. No calendar-event or binary-download reader is included.

People fetch one same-person detail per listed profile and validate the native ID, preserving externalLinks and all other returned fields. Companies likewise fetch one detail per domain. These same-resource requests add to full-read cost and require matching access.

Terminal window
bunx chkit ingest run --tag provider:circleback --tag resource:meeting_transcripts
bunx chkit ingest status --tag provider:circleback --json

Circleback exposes paginated full collections without a reliable modification feed. Every resource uses native fullSync() completion and paginate(). Each run exhausts its own provider pages; failed or interrupted reads restart from the first page, and empty successful reads still record completion. Progress never declares successful completion before destination acknowledgement. Loaded rows may remain visible after a later failure.

Link pagination preserves ownership, assignee and completion filters even on cursor-only continuations. Conflicting filters, repeated/unsafe links and malformed provider responses fail visibly. Requests forward cancellation and time out after 30 seconds; executor retries honor HTTP 429 Retry-After.

Transcripts are rechecked for every listed meeting, including unchanged old meetings. A meeting’s updatedAt is not treated as transcript readiness. The action-item listing explicitly uses assigneeType=Anyone and separately exhausts PENDING and DONE, the native statuses verified in the official OpenAPI and update contract. Unexpected statuses fail instead of silently producing a partial export.

Full reads and status partitions are mutable, not atomic snapshots. A subsequent full run reconciles late availability, edits and status changes. Absent resources remain stored, and inaccessible transcripts do not prove deletion. Full-sync readers ignore date backfill bounds and provide no date-range backfill. Use an external scheduler and one ingestion process per ClickHouse target.

Row IDs are [sourceId, nativeId] for meetings, action items and people; companies always use [sourceId, domain]. Native provider IDs and references remain unchanged in data. Keep source_id in every cross-resource join.

Company responses preserve listing and same-company detail fields. The company_id envelope metadata preserves a native list/detail ID or a consistent non-null company ID from returned people. Conflicting numeric IDs fail. Empty companies without a numeric ID remain valid domain-keyed rows with company_id: null; no numeric mapping is invented. Join person data.companyId to company company_id only when known.

Account or filter changes require a new sourceId and deliberately configured destinations. Version 0.1.0 used direct meeting JSON and unscoped keys. The 0.2.0 source envelopes, scoped IDs, string tags and additional tables require deliberate migration or reingestion; earlier rows are not automatically rewritten.

Each transcript has one stable full-content key [sourceId, meetingId] and one read-outcome key [sourceId, meetingId, 'availability']. The envelope discriminator is kind: 'transcript' or kind: 'availability'.

HTTP 200 writes the complete raw segment array and an available outcome. HTTP 403 or 404 updates only the outcome to forbidden or not_found, preserving earlier successful text. A successful empty array replaces the prior content with []. Malformed transcript data fails without marking availability successful. Segments have no synthetic IDs.

Select successful snapshots with raw.kind::String = 'transcript'; join the latest outcome separately:

WITH snapshots AS (
SELECT raw.source_id::String AS source_id,
raw.meeting_id::String AS meeting_id,
JSONExtractRaw(toJSONString(raw), 'data') AS transcript
FROM default.circleback_meeting_transcripts_raw FINAL
WHERE raw.kind::String = 'transcript'
), outcomes AS (
SELECT raw.source_id::String AS source_id,
raw.meeting_id::String AS meeting_id,
raw.status::String AS latest_status
FROM default.circleback_meeting_transcripts_raw FINAL
WHERE raw.kind::String = 'availability'
)
SELECT s.source_id, s.meeting_id, s.transcript, o.latest_status
FROM snapshots AS s
LEFT JOIN outcomes AS o
ON s.source_id = o.source_id AND s.meeting_id = o.meeting_id;

An unavailable latest outcome can coexist with a retained successful snapshot. Neither unavailable status is assumed to mean temporary processing or deletion.

Terminal window
bunx chkit add circleback --with-tests
bun test src/integrations/circleback/tests/basic.test.ts

The installed fixture suite requires no credentials or live database and verifies independently selected resources, complete status partitions, source/sink replay, late transcripts, retained content, native joins and pagination failures.

Version 0.2.0

  • Provide one configurable account pipeline with independent meetings, meeting transcripts, action items, people, and companies raw streams.
  • Use native full-sync completion and safe first-page replay with scoped identities and validated Link pagination that preserves query filters.
  • Revisit every eligible meeting for late transcripts; retain successful snapshots separately from forbidden/not-found read outcomes.
  • Read action items for anyone across both native PENDING and DONE statuses; preserve people and company relationships without inventing numeric company IDs.
  • Store complete tag names as string arrays and raw provider payloads in source-scoped envelopes; adopting this layout requires reingestion or migration of earlier unscoped meeting observations.
  • Add portable fixtures for independently selected resources, reversed stream order, malformed responses, pagination, late transcripts, and source/destination recovery.
  • Consume full pagination pages and use shared continuation-cycle detection while preserving independent resource checkpoints and provider Link validation.

Version 0.1.0

  • Introduce raw meeting ingestion with transcript enrichment.