This page cannot be shown here. You can still read it on the original site — the toolbar below keeps your place in the directory.
These are the slides for a talk I gave recently. Abstract. IPython notebooks, NumPy and Pandas data frames are the go-to tools for doing data science with Python. Spark and PySpark is rapidly becoming the de facto standard for doing analysis on large volumes of data. But what about CPU-intensive tasks? What about rough numerical, but distributed computations? In the first part of this talk I give…
Comments
Nothing yet. Say the first thing.
Sign in to join the conversation.