Posts

PySpark on Kuberntes @ Python Barcelona March Meetup

Thanks for joining me on 2019-03-21 at Python Barcelona March Meetup 2019 Barcelona, Spain for PySpark on Kuberntes. The slides are at http://bit.ly/2Fv7Uwj . PySpark on Kubernetes @ Python Barcelona March Meetup from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Contributing to Spark 3 @ Spark BCN Meetup

Thanks for joining me on 2019-03-19 at Spark BCN Meetup 2019 Barcelona, Spain for Contributing to Spark 3. The slides are at http://bit.ly/2HwlXFf . Contributing to Apache Spark 3 from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

PyData Hong Kong - Making the Big Data ecosystem work together with Python: Apache Arrow, Spark, Flink, Beam, and Dask @ PyData Hong Kong

Thanks for joining me on 2019-02-19 at PyData Hong Kong 2019 Hong Kong for PyData Hong Kong - Making the Big Data ecosystem work together with Python: Apache Arrow, Spark, Flink, Beam, and Dask.I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Validating Big Data Jobs An exploration with @ApacheSpark & @ApacheAirflow (+ friends) @ FOSDEM

Thanks for joining me on 2019-02-03 at FOSDEM 2019 Brussels, Belgium for Validating Big Data Jobs An exploration with @ApacheSpark & @ApacheAirflow (+ friends). The slides are at http://bit.ly/2wMwRiF . Validating big data pipelines - FOSDEM 2019 from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Introducing @Kubeflow (w. Special Guests Tensorflow and @ApacheSpark) @ FOSDEM

Thanks for joining us ( @holdenkarau , @rawkintrevo ) on 2019-02-03 at FOSDEM 2019 Brussels, Belgium for Introducing @Kubeflow (w. Special Guests Tensorflow and @ApacheSpark).I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Apache Spark on Kubernetes -- Avoiding the pain of YARN @ Pre-FOSDEM Belgium Kubernetes Meetup

Thanks for joining me on 2019-02-02 for Apache Spark on Kubernetes -- Avoiding the pain of YARN.The talk covered: Apache Spark is one of the most popular big data tools, and starting last year has had integrated support for running on Kubernetes. This talk will introduce some of the use cases of Apache Spark quickly (machine learning, ETL, etc.) and then look at the current cluster managers Spark runs on and their limitations. Most of the focus will be around running non-Java code, and the challenges associated with dependencies along with general challenges like scale-up & down. Its not all sunshine and roses though, I will talk about some of the limitations of our current approach and the work being done to improve this. .I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Predicting areas for PR Comments based on Code Vectors & Mailing List Data @ FOSDEM

Thanks for joining us ( @holdenkarau , @krisnova ) on 2019-02-03 at FOSDEM 2019 Brussels, Belgium for Predicting areas for PR Comments based on Code Vectors & Mailing List Data.I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Understanding Spark Tuning with Auto Tuning (or how to stop your pager going off @ Data Day Texas

Thanks for joining me on 2019-01-26 at Data Day Texas 2019 Austin, TX, USA for Understanding Spark Tuning with Auto Tuning (or how to stop your pager going off.I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Validating Big Data Jobs - Stopping Failures before Production (w/ Spark, BEAM, & friends!) -- now with Scooters @ Scala eXchange 2018

Thanks for joining me on 2018-12-14 at Scala eXchange 2018 London, UK for Validating Big Data Jobs - Stopping Failures before Production (w/ Spark, BEAM, & friends!) -- now with Scooters.I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Holden @ Kiwi Code Mania: talk title TBD @ Code Mania

Thanks for joining me on 2019-05-15 at Code Mania 2019 Auckland, New Zealand for Holden @ Kiwi Code Mania: talk title TBD.I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Validating Big Data Jobs - Stopping Failures before Production (w/ Spark, BEAM, & friends!) @ Big Data Spain

Image
Thanks for joining me on 2018-11-14 at Big Data Spain 2018 Madrid, Spain for Validating Big Data Jobs - Stopping Failures before Production (w/ Spark, BEAM, & friends!). The slides are at http://bit.ly/2S1y3HJ .The video of the talk is up at http://bit.ly/2Gxk23h . Validating Big Data Pipelines - Big Data Spain 2018 from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Big Data w/Python on Kubernetes (PySpark on K8s) @ Big Data Spain

Image
Thanks for joining me on 2018-11-15 for Big Data w/Python on Kubernetes (PySpark on K8s). The slides are at http://bit.ly/2RWsxpA .The video of the talk is up at http://bit.ly/2R9x4bE . Big data with Python on kubernetes (pyspark on k8s) - Big Data Spain 2018 from Holden Karau And some related links: recorded demos , setup livestream Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

A Basic Introduction to PySpark Dataframes by exploring ASF Gender Diversity Data - workshop @ @pyconca

Image
Thanks for joining us ( @holdenkarau , @math_foo ) on 2018-11-10 at @pyconca 2018 Toronto, ON, Canada for A Basic Introduction to PySpark Dataframes by exploring ASF Gender Diversity Data - workshop. You can find the code for this talk at https://github.com/holdenk/diversity-analytics/ .The slides are at http://bit.ly/2RNhG1g .There is a related video you might want to check out. Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

PyConCanada After Lunch Keynote Sunday @ @pyconca

Thanks for joining me on 2018-11-08 for PyConCanada After Lunch Keynote Sunday.I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Dealing With Contributor Overload @ Festival de Software Libre

Thanks for joining us ( @holdenkarau , @griscz ) on 2018-11-02 at Festival de Software Libre 2018 Puerto Vallarta, Jalisco, Mexico for Dealing With Contributor Overload.I'll update this post with the slides soon.Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Diversity in OSS with Holden & Gris @ Festival de Software Libre

Thanks for joining us ( @holdenkarau , @griscz ) on 2018-11-02 at Festival de Software Libre 2018 Puerto Vallarta, Jalisco, Mexico for Diversity in OSS with Holden & Gris. The slides are at http://bit.ly/2JBYgcw . Keynote Open Source Diversity - Festival del Software Libre from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

End to End ML with Kubeflow: Scaling with Big & Tiny Data (+ deep learning of course) @ signal

Image
Thanks for joining me on 2018-10-17 at signal 2018 San Francisco, CA, USA for End to End ML with Kubeflow: Scaling with Big & Tiny Data (+ deep learning of course). The slides are at http://bit.ly/2pWFUKj .The video of the talk is up at http://bit.ly/2GBMp0B .And if you want there is a related codelab you can try out . Intro - End to end ML with Kubeflow @ SignalConf 2018 from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Building Recoverable (and optionally Async) Spark Pipelines @ Scylla Summit 2018

Thanks for joining me on 2018-11-07 for Building Recoverable (and optionally Async) Spark Pipelines . The slides are at http://bit.ly/2D90dfy . Building Recoverable (and optionally async) Pipelines with Apache Spark (+ small revisions) from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

The magic of distributed systems: when it all breaks and why @ Reversim Summit 2018

Thanks for joining me on 2018-10-09 for The magic of distributed systems: when it all breaks and why . The slides are at http://bit.ly/2A8HfDI . The magic of (data parallel) distributed systems and where it all breaks - Reversim from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback

Validating Big Data Jobs - Stopping Failures Before Production on Apache Spark @ @sparkaisummit

Thanks for joining me on 2018-10-03 at @sparkaisummit 2018 London, UK for Validating Big Data Jobs - Stopping Failures Before Production on Apache Spark . The slides are at http://bit.ly/2QqQUea . Validating big data jobs - Spark AI Summit EU from Holden Karau Comment bellow to join in the discussion :). Talk feedback is appreciated at http://bit.ly/holdenTalkFeedback