Home

Java developer - Java developer
[email protected]
Location: Remote, Remote, USA
Relocation: bdf
Visa: d
I have used AWS Glue and PySpark to process large volumes of healthcare, equipment telemetry, operational, and financial data. Glue jobs read source data from Amazon S3, apply transformations and validations, and write the processed results back to S3 or load them into Snowflake.

To improve performance, I use partitioning, predicate pushdown, job bookmarks, appropriate Spark executor settings, and efficient file formats such as Parquet. I also control the number and size of output files to avoid the small-file problem.

Job bookmarks help us process only new or changed data instead of reprocessing the complete dataset during every execution.
Keywords: sthree

To remove this resume please click here or send an email from [email protected] to [email protected] with subject as "delete" (without inverted commas)
[email protected];7562
Enter the captcha code and we will send and email at [email protected]
with a link to edit / delete this resume
Captcha Image: