Databricks

Polytomic connects Databricks to your databases, SaaS tools, and cloud storage for ETL and reverse ETL. Load Salesforce, Postgres, Stripe, and application data into Unity Catalog tables through staged cloud storage, and sync modeled results back out to your operational tools, without writing code.

Databricks logo

Full bidirectional ETL

Sync in both directions between Databricks and your systems.

Optimized for scale

Efficient operations minimize Databricks compute costs.

Unity Catalog native

Three part catalog naming, external locations, and atomic table swaps.

Example workflows

Consolidate Salesforce, Stripe, Slack, and Outreach into Databricks

What you can sync

Move data between Databricks and your databases, SaaS tools, and cloud storage with automatic table creation, schema evolution, and Unity Catalog aware writes.

From Databricks

Tables across catalogs and schemas
Arbitrary SQL queries as model sources
Incremental reads using one or more tracking columns
STRUCT, ARRAY, MAP, and DECIMAL types
Unity Catalog discovery with three part naming

To Databricks

Schemas and tables created automatically
Incremental merge with soft or hard deletes
Create, update, create or update, replace, and append
Partitioned output tables
Optional Delta UniForm for Iceberg compatibility

How it works

Connect Databricks with a SQL warehouse and Polytomic treats it as a full lakehouse source and destination. Load from your SaaS tools, databases, and cloud storage, then read modeled results back out to your CRM and operational systems on the same platform.

Writes stage through your own cloud storage and load with COPY INTO. One consequence worth knowing is that data becomes queryable in Databricks while a large initial sync is still working through its backlog, rather than appearing only when the whole job finishes.

  • Connect with a personal access token or an OAuth service principal
  • Configure cloud storage staging for destination writes
  • Discover catalogs, schemas, and tables under Unity Catalog
  • Read in full, incrementally on tracking columns, or by SQL query
  • Create schemas and tables and add columns as your source evolves
  • Monitor sync health, record volume, and failures from one interface

How to set up

Polytomic authenticates to Databricks with a personal access token or an OAuth service principal using client credentials. You supply the workspace hostname and the HTTP path of a SQL warehouse. Source only connections need nothing more.

How to get connected
  1. 1In Databricks, note your workspace hostname and SQL warehouse HTTP path
  2. 2Create a personal access token, or an OAuth service principal
  3. 3In Polytomic, navigate to Connections, then Add Connection, then Databricks
  4. 4Enter the hostname, HTTP path, and credentials
  5. 5For destination use, add S3 or ADLS Gen2 staging credentials
  6. 6Test the connection and click Save

See the documentation for full setup details.

Frequently asked questions

Why teams choose Polytomic

No engineering required

Set up and manage syncs without writing code.

Flexible data modeling

Use SQL to define exactly what data gets synced.

Handles scale automatically

Supports large datasets with incremental syncs and bulk APIs.

Start syncing your data with Databricks in minutes

No credit card required. Free trial available.

More integrations beyond Databricks