Databricks Writer programmer's reference
Databricks Writer properties
property | type | default value | notes |
|---|---|---|---|
Authentication Type | enum | Personal Access Token | Appears in Flow Designer only when Use Connection Profile is False. With the default setting PersonalAccessToken, Striim's connection to Databricks is authenticated using the token specified in Personal Access Token. Alternatively, select Manual OAuth or Service Principal. For more information, see Initial setup for Databricks Writer. |
CDDL Action | enum | Process | See Handling schema evolution. If TRUNCATE commands may be entered in the source and you do not want to delete events in the target, precede the writer with a CQ with the select statement |
Client ID | string | Appears in Flow Designer only when Connection Profile is False and Authentication Type is Manual OAuth. This property is required when Manual OAuth is selected as the value of the Authentication Type property. | |
Client Secret | encrypted password | Appears in Flow Designer only when Connection Profile is False and Authentication Type is Manual OAuth. This property is required when Manual OAuth is selected as the value of the Authentication Type property. | |
Connection Profile Name | Enum | Appears in Flow Designer only when Use Connection Profile is True. See Connection profiles. | |
Connection Retry Policy | String | initialRetryDelay=10s, retryDelayMultiplier=2, maxRetryDelay=1m, maxAttempts=5, totalTimeout=10m | Do not change unless instructed to by Striim support. |
Connection URL | String | Appears in Flow Designer only when Use Connection Profile is False. Provide the JDBC URL from the JDBC/ODBC tab of the Databricks cluster's Advanced options (see Get connection details for a cluster). If the URL starts with | |
External Stage Connection Profile Name | enum | Appears in Flow Designer only when Use Connection Profile is True and External Stage Type is ADLSGen2, GCS, or S3. Select or specify the name of the connection profile for the external stage. (When Databricks Writer uses a connection profile, you must use a connection profile for ADLSGen2 or S3 as well.) | |
External Stage Type | enum |
| Set to ADLSGen2, GCS, or S3 to match the stage type you chose in Initial setup for Databricks Writer. NotePersonal staging locations have been deprecated by AWS (see Create metastore-level storage) and Microsoft (see Create metastore-level storage). |
Ignorable Exception Code | String | Set to TABLE_NOT_FOUND to prevent the application from terminating when Striim tries to write to a table that does not exist in the target. See Handling "table not found" errors for more information. Ignored exceptions will be written to the application's exception store (see CREATE EXCEPTIONSTORE). | |
Mode | enum | Append Only | Set to Merge if that was your choice in Building pipelines with Databricks Writer. |
Optimized Merge | Boolean | false | Appears in Flow Designer only when Mode is Merge. Set to True only when Mode is MERGE and the target's input stream is the output of an HP NonStop reader, MySQL Reader, or Oracle Reader source and the source events will include partial records. For example, with Oracle Reader, when supplemental logging has not been enabled for all columns, partial records are sent for updates. When the source events will always include full records, leave this set to False. When Optimized Merge is True, each primary key update is handled in a separate write operation. If the source has frequent primary key updates, this may lead to a decline in write performance compared with Optimized Merge = False. |
Parallel Threads | Integer | Not supported when Mode is Merge. | |
Personal Access Token | encrypted password | Appears in Flow Designer only when Connection Profile is False and Authentication Type is Personal Access Token. Used to authenticate with the Databricks cluster (see Generate a personal access token). The user associated with the token must have read and write access to DBFS (see Important information about DBFS permissions). If table access control has been enabled, the user must also have MODIFY and READ_METADATA (see Data object privileges - Data governance model). | |
Personal Staging User Name | String | Personal staging locations have been deprecated by AWS (see Create metastore-level storage) and Microsoft (see Create metastore-level storage). | |
Refresh Token | encrypted password | Appears in Flow Designer only when Connection Profile is False and Authentication Type is Manual OAuth. This property is required when Manual OAuth is selected as the value of the Authentication Type property. The token expires in 90 days, after which the application will halt. To avoid that, use a connection profile (see Connection profiles), which will allow you to update the token without stopping the application. Alternatively, prior to expiry stop the application and update the token. | |
Stage Location | String |
| If you choose DBSROOT as your External Stage Location (not recommended), set the path to the staging area in DBFS here, for example, |
Tables | String | The name(s) of the table(s) to write to. If not using a Database Reader source with Create Schema enabled or a wizard with initial schema creation, the table(s) must already exist in the database. Specify target table names as When the target's input stream is a user-defined event, specify a single table. The only special character allowed in target table names is underscore ( When the input stream of the target is the output of a DatabaseReader, IncrementalBatchReader, or SQL CDC source (that is, when replicating data from one database to another), it can write to multiple tables. In this case, specify the names of both the source and target tables. You may use the source.emp,target_database.emp source_schema.%,target_catalog.target_database.% source_database.source_schema.%,target_database.% source_database.source_schema.%, target_catalog.target_database.% MySQL and Oracle names are case-sensitive, SQL Server names are not. Specify names as If a column in the target table is not mapped to a column in the source (see When is a target column not mapped to a source column?):
See Mapping columns and Defining relations between source and target using ColumnMap and KeyColumns for additional options. | |
Tenant ID | String | Appears in Flow Designer only when Connection Profile is False and Authentication Type is Manual OAuth. This property is required when Manual OAuth is selected as the value of the Authentication Type property. | |
Upload Policy | String | eventcount:100000, interval:60s | The upload policy may include eventcount and/or interval (see Setting output names and rollover / upload policies for syntax). Buffered data is written to the external stage every time any of the specified values is exceeded. With the default value, data will be written every 60 seconds or sooner if the buffer contains 100,000 events. When the app is quiesced, any data remaining in the buffer is written to the storage account; when the app is undeployed, any data remaining in the buffer is discarded. |
Use Connection Profile | Boolean | False | Set to True to use a connection profile instead of specifying the connection properties in the adapter properties. See Connection profiles. |
Configuration by Cloud Endpoint
Writing to Azure Databricks
Concept: Azure Databricks is the only endpoint that supports Microsoft Entra ID authentication in addition to Personal Access Token and Service Principal (M2M OAuth). Staging uses Azure Data Lake Storage Gen2.
Azure Data Lake Storage (ADLS) Gen2 properties for Databricks Writer
property | type | default value | notes |
|---|---|---|---|
Azure Account Access Key | encrypted password | When Authentication Type is set to ServiceAccountKey, specify the account access key from Storage accounts > <account name> > Access keys. When Authentication Type is set to AzureAD, this property is ignored in TQL and not displayed in the Flow Designer. | |
Azure Account Name | String | the name of the Azure storage account for the blob container | |
Azure Container Name | String | striim-deltalakewriter-container | the blob container name from Storage accounts > <account name> > Containers If it does not exist, it will be created. |
Example (Entra ID authentication with ADLS Gen2 staging):
Using EntraID as the authentication type requires a Connection Profile.
CREATE OR REPLACE TARGET db USING Global.DeltaLakeWriter ( useConnectionProfile: true, connectionProfileName: 'admin.DatabricksEntraIDCP', stageLocation: '/', CDDLAction: 'Process', ConnectionRetryPolicy: 'initialRetryDelay=10s, retryDelayMultiplier=2, maxRetryDelay=1m, maxAttempts=5, totalTimeout=10m', Mode: 'APPENDONLY', externalStageType: 'ADLSGen2', Tables: 'public.sample_pk,sample_catalog.sample_db.sample_pk', externalStageConnectionProfileName: admin.ADLSGen2CP'', uploadPolicy: 'eventcount:10000,interval:60s' ) INPUT FROM DBOut;
Operational considerations: to use an ADLS Gen2 container as your staging area, your Databricks instance should run Databricks Runtime 11.0 or later (10.4 or later at minimum for any external stage). When Authentication Type is Entra ID, the Azure Account Access Key property is ignored — authentication for staging is handled through the same Entra ID credential. We recommend using a Connection Profile for Entra ID rather than configuring Manual OAuth directly, since a connection profile is significantly simpler and lets you rotate credentials without stopping the application.
Writing to Databricks on AWS
Concept: Databricks on AWS supports Personal Access Token, Manual OAuth, and Service Principal (M2M OAuth) authentication (no Entra ID — that's Azure-specific). Staging uses Amazon S3.
Amazon S3 properties for Databricks Writer
property | type | default value | notes |
|---|---|---|---|
S3 Access Key | String | an AWS access key ID (created on the AWS Security Credentials page) for a user with read and write permissions on the bucket If the Striim host has default credentials stored in the | |
S3 Bucket Name | String | striim-deltalake-bucket | Specify the S3 bucket to be used for staging. If it does not exist, it will be created. |
S3 Region | String | us-west-1 | the AWS region of the bucket |
S3 Secret Access Key | encrypted password | the secret access key for the access key If the Striim host has default credentials stored in the |
Operational considerations: to use an S3 bucket as your staging area, your Databricks instance should run Databricks Runtime 11.0 or later. See Setup for Databricks on AWS for creating the bucket and its IAM policy.
Writing to Databricks on Google Cloud
Concept: Databricks on Google Cloud supports Personal Access Token, Manual OAuth, and Service Principal (M2M OAuth) authentication (no Entra ID — that's Azure-specific). Staging uses Google Cloud Storage, which — unlike S3 and ADLS Gen2 — requires a Storage Credential prerequisite in Unity Catalog because Databricks' COPY INTO ... WITH CREDENTIAL clause does not accept GCS authentication parameters directly. See Setup for Databricks on Google Cloud.
Google Cloud Storage (GCS) properties for Databricks Writer
property | type | default value | notes |
|---|---|---|---|
Project ID | String |
| Specify the Google Cloud project ID of the Google Cloud Storage bucket. |
Service Account Key | String |
| Upload the service account key you downloaded as instructed in Create the service account and download the service account key. Alternatively, you can leave this field blank and set the key path using the GOOGLE_APPLICATION_CREDENTIALS environment variable. Restart Striim after setting the environment variable. |
Connection Retry | String |
| Optionally adjust this setting (see GCS Writer properties). |
Example — with a GCS Connection Profile:
CREATE OR REPLACE APPLICATION DatabricksDBR; CREATE OR REPLACE SOURCE DBSource USING DatabaseReader ( Tables: '"TPCC"."H_ORDER"', QuiesceOnILCompletion: true, adapterName: 'DatabaseReader', connectionProfileName: '', DatabaseProviderType: 'MySQL', ParallelThreads: 1, RestartBehaviourOnILInterruption: 'keepTargetTableData', useConnectionProfile: false, ConnectionURL: 'jdbc:mysql://localhost:3306?allowPublicKeyRetrieval=true&rewriteBatchedStatements=true&disableMariaDbDriver&useSSL=false', FetchSize: 10000, CreateSchema: true, Username: 'root' ) OUTPUT TO DBOut; CREATE OR REPLACE TARGET DatabricksTarget USING DeltaLakeWriter ( optimizedMerge: false, stageLocation: '/', useConnectionProfile: true, CDDLAction: 'Process', adapterName: 'DeltaLakeWriter', ConnectionRetryPolicy: 'initialRetryDelay=10s, retryDelayMultiplier=2, maxRetryDelay=1m, maxAttempts=5, totalTimeout=10m', Mode: 'APPENDONLY', connectionProfileName: 'admin.Databricks_Connection_Profile', Tables: 'TPCC.H_ORDER,my_catalog.tpcc.h_order', externalStageType: 'GCS', externalStageConnectionProfileName: 'admin.gcscp1', storageCredentialName: 'gcsstagecredential', gcsBucketName: 'striim-deltalake-bucket', gcsBucketRegion: 'asia-south1', uploadPolicy: 'eventcount:100000,interval:60s', gcsProjectId: '' ) INPUT FROM DBOut; END APPLICATION DatabricksDBR;
Operational considerations: if you want Databricks Writer to automatically create the GCS bucket when it doesn't exist, the Google Cloud service account backing your Storage Credential must have the storage.buckets.create permission in addition to the standard read/write permissions. GCS is currently supported only for Databricks hosted on Google Cloud — it cannot be used as a staging area for Databricks on AWS or Azure.
Databricks Writer Data Type Support and Correspondence
TQL type | Delta Lake type |
|---|---|
java.lang.Byte | binary |
java.lang.Double | double |
java.lang.Float | float |
java.lang.Integer | int |
java.lang.Long | bigint |
java.lang.Short | smallint |
java.lang.String | string |
org.joda.time.DateTime | timestamp |
For additional data type mappings, see Data type support & mapping for schema conversion & evolution.
Writing to Apache Iceberg tables in Databricks
Databricks Writer can write to existing Apache Iceberg tables in Databricks. No special Databricks Writer configuration is required. The tables must be created and configured as described in Read Delta tables with Iceberg clients. Iceberg tables cannot be created using a Striim initial load wizard and schema evolution (CDDL) is not supported.
The following table shows how Databricks Writer maps Databricks data types to Iceberg data types. Only Iceberg specification versions 1 and 2 are supported.
Databricks data type | Iceberg data type |
|---|---|
ARRAY<Type> | list<Type> |
BIGINT | long |
BINARY | uuid |
BINARY | fixed(L) |
BINARY | binary |
BOOLEAN | boolean |
DATE | date |
DECIMAL(p/s) | decimal(p/s) |
DOUBLE | double |
FLOAT | float |
INT | int |
MAP | not supported |
STRING | string |
STRUCT | not supported |
TIME | time |
TIMESTAMP | timestamptz |
TIMESTAMP_NTZ | not supported |