Skip to main content

GCP Lakehouse Writer operational considerations

Use these sections to monitor, troubleshoot, tune, and validate GCP Lakehouse Writer applications after deployment. For setup-time Spark sizing decisions, also review Performance guidance and best practices before you create or resize the cluster.

Monitoring

GCP Lakehouse Writer exposes table-level and adapter-level metrics in Striim. Use these metrics with Google Cloud cluster and job monitoring to understand throughput, latency, queued work, and failures.

Table write information

Table Write Info is a JSON array that lists table-level metrics.

Sub-metric

Description

Frequency

Mapped Source Table

Source table mapped to the target table.

Per batch

Last batch info

Metrics from the last executed batch.

Per batch

Last successful merge time

Time of the last executed task for a table.

Per batch

Last Applied DDL Time

Time of the last DDL executed for a table.

Per DDL batch

Last Applied DDL Statement

Last DDL statement executed for a table.

Per DDL batch

Total Batches Created

Total tasks created for a table.

Per batch

Total Batches Queued

Tasks currently queued for a table.

Per batch

Total Batches Ignored

Tasks ignored for a table.

Per batch

Total Batches Uploaded

Tasks successfully executed for a table.

Per batch

Total event info

Overall event count for a table.

Per batch

Avg Upload Time in ms

Average upload time across batches.

Per batch

Avg Compaction Time in ms

Average compaction time across batches.

Per batch

Avg In-Mem Compaction Time in ms

Average in-memory compaction time across batches.

Per batch

Avg Merge Time in ms

Average merge time across batches.

Per batch

Avg Waiting Time in Queue in ms

Average time batches spent waiting in the queue.

Per batch

Avg Event Count Per Batch

Average event count in a batch.

Per batch

Avg Batch Size in bytes

Average batch size.

Per batch

Avg Stage Resources Management Time in ms

Average time to clear or create staging resources, such as stage table and staging area.

Per batch

Avg Integration Time in ms

Average time to move a processed batch to the target table.

Per batch

Min Integration Time in ms

Minimum time for a processed batch to reach the target table.

Per batch

Max Integration Time in ms

Maximum time for a processed batch to reach the target table.

Per batch

Adapter-level metrics

Metric

Description

Frequency

Write Timestamp

Last time a batch was executed across all tables.

Per batch

Target Freshness

Time since a batch was executed across all tables.

Per batch

Discarded Event Count

Total events discarded across all tables.

Per event / batch

Connection Retry Information

Total reconnects and last known reconnect time.

When available

Queued Batches Size In Bytes

Size of all queued batches.

When available

Monitor GCP Managed Apache Spark

Use Google Cloud monitoring tools to investigate cluster and job behavior:

  • Open the Managed Service for Apache Spark cluster page and review cluster metrics such as memory, HDFS capacity, network traffic, CPU utilization, and disk operations.

  • Open Web Interfaces to access Spark History Server and YARN Resource Manager.

  • Open Jobs to view Managed Service for Apache Spark jobs. Click a job to inspect configuration, status, monitoring data, and logs.

  • Use the Spark History Server to view job status, logs, stages, executor assignment, and execution timeline.

  • Use YARN Resource Manager to view node status, health, available memory, and allocated applications.

When a Striim application halts because of a Spark job failure, use the job ID reported by Striim or the SHOW command to find the corresponding job in Google Cloud.

Performance guidance and best practices

Cluster sizing

GCP Managed Apache Spark clusters require a master node and at least one worker node. The n2-standard-4 VM type is a minimum-cost option capable of running Iceberg batches, but larger workloads may require larger VMs or more workers.

You cannot change individual VM configuration after cluster creation, so choose the master and worker machine sizes before creating the cluster. You can increase or decrease the number of worker nodes later based on workload.

Workers and parallelism

Spark performance can be increased by increasing parallelism. You can do this through:

  • vertical scaling: use VMs with more vCPUs, memory, and disk space

  • horizontal scaling: increase the number of worker nodes

Secondary workers can be added to a Dataproc cluster, but their availability is not guaranteed.

Spark configuration

Managed Service for Apache Spark determines the number of executors based on the number of workers and vCPUs per worker. You can override executor count, cores, memory, driver memory, and shuffle partitions with AdditionalConfiguration in the GCP Managed Apache Spark Connection Profile.

GCP Lakehouse Writer sets the following Spark enhancement properties to true for jobs it creates:

  • spark.dataproc.enhanced.optimizer.enabled

  • spark.dataproc.enhanced.execution.enabled

Example scenarios

Scenario 1: 1000 source tables and 8 writer instances

If a source database contains 1000 tables and 8 GCP Lakehouse Writer instances split the tables evenly, each writer handles about 125 tables. Because all 8 writers can submit batches simultaneously, the cluster should have at least 8 executor cores for task allocation.

A single worker with 8 vCPUs or two workers with 4 vCPUs each provide the same number of worker cores. However, a single n2-standard-8 worker can be preferable to two n2-standard-4 workers because less memory is consumed by Spark/system overhead and more cumulative memory is available to executors.

Scenario 2: One highly active table

For one high-traffic table, such as orders, you can optimize by using Spark scheduling and separating the high-traffic table into its own writer instance:

  1. Create a separate Spark scheduling policy and restart the cluster.

  2. Create one writer for the high-traffic table and another writer for other tables.

  3. Create a GCP Managed Apache Spark Connection Profile for the high-traffic writer and add the scheduling policy in additional configuration.

  4. Assign the new Connection Profile to the high-traffic writer.

  5. Run both writers.

This approach can improve priority for the high-traffic table, but it can starve jobs from the other writer.

Scenario 3: Initial-load parallelism

The Spark cluster should have at least as many worker cores as the number of ParallelThreads. For example, if ParallelThreads is 8, the cluster should have either one worker with 8 vCPUs or two workers with 4 vCPUs each. More cores and memory can improve performance.

Scenario 4: Batch size

A higher batch event count can improve performance. In observed testing, a batch with 1 million operations performed up to approximately 3 times faster than a batch with 100,000 operations. Treat this as workload-dependent guidance, not a guarantee.

Before increasing eventcount, consider:

  • whether Spark executor memory is sufficient for the larger batch

  • whether other GCP Lakehouse Writer instances share the same cluster memory

  • whether the larger batch file increases upload time enough to affect latency

Limitations

  • In REST catalog mode, IcebergTablesLocation must use a root-level bucket path such as gs://my-bucket or gs://my-bucket/. Subdirectory paths such as gs://my-bucket/subdirectory are not supported.

  • Spark does not support the Iceberg TIME type natively. GCP Lakehouse Writer handles it as string.

  • GCP Lakehouse Writer does not automatically create partitioned tables. Create partitioned Iceberg tables outside Striim before writing to them.

  • In MERGE mode, the only supported special character in table names is underscore (_).

  • Some Managed Service for Apache Spark job failures require inspection of Dataproc and Spark logs in Google Cloud to identify the exact root cause.