| Java developer - Java developer |
| [email protected] |
| Location: Remote, Remote, USA |
| Relocation: bdf |
| Visa: d |
|
I have used AWS Glue and PySpark to process large volumes of healthcare, equipment telemetry, operational, and financial data. Glue jobs read source data from Amazon S3, apply transformations and validations, and write the processed results back to S3 or load them into Snowflake.
To improve performance, I use partitioning, predicate pushdown, job bookmarks, appropriate Spark executor settings, and efficient file formats such as Parquet. I also control the number and size of output files to avoid the small-file problem. Job bookmarks help us process only new or changed data instead of reprocessing the complete dataset during every execution. Keywords: sthree |