I've written a book "Fast Data Processing with Spark" which covers Python, Scala, and Java

I recently finished writting "Fast Data Processing with Spark". Apache Spark is a framework for writing fast, distributed programs.


Fast Data Processing with Spark covers how to write distributed map reduce style programs with Spark. The book guides you through every step required to write effective distributed programs from setting up your cluster and interactively exploring the API, to deploying your job to the cluster.


Personally, while the fast nature of Spark is not to be understated, I really enjoy its functional style APIs and find it a lovely environment to code in.

Comments

Popular posts from this blog

New year new job, same projects

Making Hibernate work on Ubuntu 22.04 (jammy) on the Framework Laptop w/full disk encryption

Unexpected gmail account name disclosure on calendars