gcp-spark
Develops, optimizes and executes Spark code on Managed Spark on Google Cloud (Dataproc Clusters and Serverless). Reads and writes data using BigLake Iceberg catalogs, BigQuery and Spanner. Debugs execution failures. Use when: - Writing Spark ETL pipelines on Google Cloud Platform. - Optimizing PySpark or Spark SQL code for performance, memory, or OOM risks. - Preparing Spark workloads for production submission. - Training or running inference with Machine Learning models with spark on Google Cloud Platform. - Managing Spark clusters, jobs, batches, and interactive sessions. Don't use when: - Writing generic Python scripts that don't use Spark. - Performing simple SQL queries that can be done directly in BigQuery. - Troubleshooting failed Spark workloads or analyzing logs (use @skill:gcp-spark-troubleshooting).
- Version
- v20
- License
- Apache-2.0
Pinned to revision 0dd8034c50d6, so it is the text this page describes rather than whatever the author pushed since.
Files
- skills/gcp-spark/SKILL.md
- skills/gcp-spark/references/gcloud_dataproc.md
- skills/gcp-spark/references/ml_tasks.md
- skills/gcp-spark/references/read_write_data.md
- skills/gcp-spark/references/schema_direct_inspection.md
- skills/gcp-spark/references/spark_optimizations.md
- skills/gcp-spark/references/spark_refactoring_guide.md
Every link opens the file at its source, pinned to the revision this page describes.