Troubleshooting Databricks Writer
The table below describes the general cause of common runtime issues. Since the exact wording of error messages can change between Striim releases, this table describes the nature and cause of each issue rather than an exact error string — check your application's logs and exception store for the specific message.
Symptom | Likely cause |
|---|---|
Application halts shortly after deployment with an authentication-related error | Invalid or expired Personal Access Token, Client ID, or Client Secret; or a Connection URL that doesn't match your Databricks workspace |
Application halts with a connection-related error | The Databricks SQL warehouse or cluster is unreachable, stopped, or deleted |
Application halts with a permission-related error during CDC (after initial load succeeded) | The connecting identity (PAT user, Service Principal, or Entra ID identity) lost MODIFY privilege on the target tables |
Application halts with a permission-related error referencing the catalog or schema | The connecting identity lost USE CATALOG or USE SCHEMA privilege |
COPY INTO fails with a permission-related error (GCS staging only) | The connecting identity lost READ FILES permission on the Storage Credential, or the GCS service account lost object-level permissions on the bucket |
GCS bucket creation fails with a permission-related error | The GCS service account lacks the storage.buckets.create permission |
Application halts referencing an expired token (Manual OAuth only) | The Refresh Token has passed its 90-day expiration — use a Connection Profile instead of Manual OAuth to avoid this |
Troubleshooting Databricks Writer's Upload Policy
Growing Upload Backlog (Total Batches Queued Trending Up)
Likely causes:
Event count set too high with no interval backstop — batches take a long time to fill, so uploads run infrequently but each run is large and slow.
Interval set lower than the time it actually takes to upload a batch — new batches queue faster than existing ones can drain.
The Databricks SQL warehouse is slow or undersized — many small statements from a too-low interval add per-statement overhead, or the warehouse is auto-suspended.
What to check:
Watch the Total Batches Queued metric. If it's climbing steadily, size Upload Policy's event count to your realistic burst volume, and don't set the interval lower than your observed end-to-end flush time.
Backpressure (Ingestion Stalls, Memory-Related Log Messages)
Likely causes:
Buffered in-memory data across all tables in the writer component exceeded available memory, most often during Optimized Merge.
Many tables sharing one Databricks Writer component, or a traffic burst that outpaces the flush rate.
What to check:
If this happens under normal (non-burst) traffic, your event count and interval are letting batches grow too large before they flush — tune Upload Policy first. If it persists, contact your Striim field engineer or Striim Support for guidance on memory tuning for Optimized Merge.
Slow Uploads or High Latency Before Data Appears in the Target Table
Likely causes:
Event count set too high with no interval — data doesn't flush until the (large) count is reached.
The staged file has grown large before upload, increasing payload size and upload time.
The underlying Databricks SQL warehouse is slow, cold-started, or undersized.
Retry loops on a failing statement add latency before succeeding or exhausting retries.
What to check:
Compare event count against your actual burst volume, confirm interval isn't set unrealistically high, and check the Databricks SQL warehouse's sizing and auto-scaling settings — once a statement is dispatched, its execution time is outside Databricks Writer's control.