Mastering Machine Learning with Spark 2.x by Michal Malohlava
Privacy Policy
Read using
(price excluding SST)
Author:
Michal Malohlava
Category:
Engineering & IT
ISBN:
9781785282416
Publisher:
Packt Publishing
File Size:
15.03 MB
(price excluding SST)
Synopsis
Key FeaturesProcess and analyze big data in a distributed and scalable wayWrite sophisticated Spark pipelines that incorporate elaborate extractionBuild and use regression models to predict flight delays Book DescriptionThe purpose of machine learning is to build systems that learn from data. Being able to understand trends and patterns in complex data is critical to success; it is one of the key strategies to unlock growth in the challenging contemporary marketplace today. With the meteoric rise of machine learning, developers are now keen on finding out how can they make their Spark applications smarter.This book gives you access to transform data into actionable knowledge. The book commences by defining machine learning primitives by the MLlib and H2O libraries. You will learn how to use Binary classification to detect the Higgs Boson particle in the huge amount of data produced by CERN particle collider and classify daily health activities using ensemble Methods for Multi-Class Classification.Next, you will solve a typical regression problem involving flight delay predictions and write sophisticated Spark pipelines. You will analyze Twitter data with help of the doc2vec algorithm and K-means clustering. Finally, you will build different pattern mining models using MLlib, perform complex manipulation of DataFrames using Spark and Spark SQL, and deploy your app in a Spark streaming environment.What you will learnUse Spark streams to cluster tweets onlineRun the PageRank algorithm to compute user influencePerform complex manipulation of DataFrames using SparkDefine Spark pipelines to compose individual data transformationsUtilize generated models for off-line/on-line predictionTransfer the learning from an ensemble to a simpler Neural NetworkUnderstand basic graph properties and important graph operationsUse GraphFrames, an extension of DataFrames to graphs, to study graphs using an elegant query languageUse K-means algorithm to cluster movie reviews datasetAbout the AuthorAlex Tellez is a life-long data hacker/enthusiast with a passion for data science and its application to business problems. He has a wealth of experience working across multiple industries, including banking, health care, online dating, human resources, and online gaming. Alex has also given multiple talks at various AI/machine learning conferences, in addition to lectures at universities about neural networks. When hes not neck-deep in a textbook, Alex enjoys spending time with family, riding bikes, and utilizing machine learning to feed his French wine curiosity!Max Pumperla is a data scientist and engineer specializing in deep learning and its applications. He currently works as a deep learning engineer at Skymind and is a co-founder of aetros.com. Max is the author and maintainer of several Python packages, including elephas, a distributed deep learning library using Spark. His open source footprint includes contributions to many popular machine learning libraries, such as keras, deeplearning4j, and hyperopt. He holds a PhD in algebraic geometry from the University of Hamburg.Michal Malohlava, creator of Sparkling Water, is a geek and the developer; Java, Linux, programming languages enthusiast who has been developing software for over 10 years. He obtained his PhD from Charles University in Prague in 2012, and post doctorate from Purdue University.During his studies, he was interested in the construction of not only distributed but also embedded and real-time, component-based systems, using model-driven methods and domain-specific languages. He participated in the design and development of various systems, including SOFA and Fractal component systems and the jPapabench control system.Now, his main interest is big data computation. He participates in the development of the H2O platform for advanced big data math and computation, and its embedding into Spark engine, published as a project called Sparkling Water.Table of ContentsIntroduction to Large Scale Machine LearningDetecting Dark Matter: The Higgs-Boson ParticleEnsemble Methods for Multi-Class ClassificationPredicting Movie Reviews using NLP and Spark StreamingOnline Learning with Word2VecExtracting Patterns from Clickstream DataGraph Analytics with GraphXLending Club Loan Prediction
Reviews
Be the first to review this e-book.
Write your review
Wanna review this e-book? Please Sign in to start your review.