Troubleshooting GCP Lakehouse Writer
Unless otherwise noted, the following runtime conditions halt the application. If skip-on-failure is enabled, Striim deactivates the affected target table and discards the offending batch instead of halting the whole application.
DDL and schema failures
Condition | Cause | Message or symptom |
|---|---|---|
Unsupported DDL received from source | Source sent RENAME TABLE, RENAME COLUMN, primary-key DDL, or constraint DDL. | Unsupported DDL operation or The DDL Operation performed is not supported |
DROP COLUMN on a mapped column | Dropped column is explicitly referenced in ColumnMap. | DROP COLUMN is not supported on columns that are explicitly mentioned in the column mapper... |
Key column missing or unmapped | Configured keycolumns entry does not exist or is not mapped. | Key-column-not-present or not-mapped error. |
Schema not found | Namespace referenced in Tables does not exist. | Schema-not-found error. |
Target table not found | Table referenced by an incoming DDL or data event cannot be found in GCP Lakehouse Runtime Catalog. | Table {<name>} not found |
Storage and data lake connectivity
Condition | Message |
|---|---|
External stage GCS bucket does not exist or is not reachable. | Unable to connect to the External Stage Location. |
Iceberg tables GCS bucket does not exist or is not reachable. | Unable to connect to the Iceberg Tables Location. |
Folder path inside the bucket does not exist. | File-not-found error for the configured path. |
GCS service account key is missing, unreadable, or invalid. | Service-account-key error during Connection Profile validation. |
GCS access is unauthorized for the configured account. | Unauthorized-access error. |
Project and configuration mismatches
Condition | Message or resolution |
|---|---|
External staging GCS project ID differs from the GCP Managed Apache Spark compute project ID. | Google project ID mismatch between external staging location and GCP Managed Apache Spark. |
GCP Lakehouse Runtime Catalog project ID differs from the data lake project ID. | Verify that the GCP Lakehouse Runtime Catalog Connection Profile and GCS data lake Connection Profile point to the intended Google Cloud projects. If your deployment uses separate projects, verify the cross-project IAM configuration before restarting. |
Compute job failures
Condition | Message or symptom |
|---|---|
Spark session could not be created. | Spark session creation failed. |
Spark ran out of cluster memory or cores. | The Spark execution was halted due to a runtime issue with the cluster |
Catalog is not reachable from Spark. | Catalog-not-reachable error. |
Catalog authentication failed after retries. | Catalog authentication error. |
Invalid catalog properties or permissions. | Catalog/permission-invalid error. |
Spark cannot write to the external staging area. | Staging-area-invalid error. |
Compute cluster did not finish pre-run checks in time. | Pre-run-check-not-completed error. |
Managed Service for Apache Spark job failed. | Root cause is populated from the job status details. |
Managed Service for Apache Spark job was cancelled. | The Striim application has halted because the compute engine Spark job for this adapter is cancelled. |
Managed Service for Apache Spark job did not start in time after retries. | The Striim application has halted because the compute job did not start within the expected time after exhausting all retry attempts. |
GCP Managed Apache Spark service account key is invalid, unreadable, or lacks privileges. | Managed Service for Apache Spark connection-profile authentication error. |
Data quality failures
Condition | Message |
|---|---|
A NOT NULL column in the target Iceberg table receives a NULL value. | A NOT NULL column in the target Iceberg table received a NULL value |
A NOT NULL target column is not mapped to any source column in MERGE mode. | A NOT NULL target column in the Iceberg table is not mapped to any source column |
Common Spark errors
Error message | Likely cause | Resolution |
|---|---|---|
Task was not acquired | Out-of-memory or memory pressure prevents the master node from acquiring the task. Multiple large jobs running simultaneously can cause this. | Increase worker nodes, use larger VMs, or reduce concurrent large jobs. |
No agent found to be active | Out-of-memory or memory pressure makes the server or cluster unhealthy. | Stop and restart the cluster, then review concurrency and memory use. |
Task not found | Cluster was deleted or stopped while a job was running. | Let jobs complete before stopping the cluster. |
Driver received SIGTERM/SIGKILL signal and exited with 143 code | Spark driver on the master node ran out of memory. | Use a master node with more memory or reduce job memory pressure. |
For job delays and VM out-of-memory scenarios, review the Managed Service for Apache Spark / Dataproc documentation and the Spark job logs in Google Cloud.