Mastering Spark for Data Science by Matthew Hallett

Mastering Spark for Data Science by Matthew Hallett from  in  category
Privacy Policy
Read using
(price excluding SST)
Author: Matthew Hallett
Category: Engineering & IT
ISBN: 9781785888281
File Size: 44.00 MB
Format: EPUB (e-book)
DRM: Applied (Requires eSentral Reader App)
(price excluding SST)

Synopsis

Key FeaturesDevelop and apply advanced analytical techniques with SparkLearn how to tell a compelling story with data science using Sparks ecosystemExplore data at scale and work with cutting edge data science methodsBook DescriptionData science seeks to transform the world using data, and this is typically achieved through disrupting and changing real processes in real industries. In order to operate at this level you need to build data science solutions of substance –solutions that solve real problems. Spark has emerged as the big data platform of choice for data scientists due to its speed, scalability, and easy-to-use APIs.This book deep dives into using Spark to deliver production-grade data science solutions. This process is demonstrated by exploring the construction of a sophisticated global news analysis service that uses Spark to generate continuous geopolitical and current affairs insights.You will learn all about the core Spark APIs and take a comprehensive tour of advanced libraries, including Spark SQL, Spark Streaming, MLlib, and more.You will be introduced to advanced techniques and methods that will help you to construct commercial-grade data products. Focusing on a sequence of tutorials that deliver a working news intelligence service, you will learn about advanced Spark architectures, how to work with geographic data in Spark, and how to tune Spark algorithms so they scale linearly.What you will learnLearn the design patterns that integrate Spark into industrialized data science pipelinesSee how commercial data scientists design scalable code and reusable code for data science servicesExplore cutting edge data science methods so that you can study trends and causalityDiscover advanced programming techniques using RDD and the DataFrame and Dataset APIsFind out how Spark can be used as a universal ingestion engine tool and as a web scraperPractice the implementation of advanced topics in graph processing, such as community detection and contact chainingGet to know the best practices when performing Extended Exploratory Data Analysis, commonly used in commercial data science teamsStudy advanced Spark concepts, solution design patterns, and integration architecturesDemonstrate powerful data science pipelinesAbout the AuthorAndrew Morgan is a specialist in data strategy and its execution, and has deep experience in the supporting technologies, system architecture, and data science that bring it to life. With over 20 years of experience in the data industry, he has worked designing systems for some of its most prestigious players and their global clients – often on large, complex and international projects. In 2013, he founded ByteSumo Ltd, a data science and big data engineering consultancy, and he now works with clients in Europe and the USA. Andrew is an active data scientist, and the inventor of the TrendCalculus algorithm. It was developed as part of his ongoing research project investigating long-range predictions based on machine learning the patterns found in drifting cultural, geopolitical and economic trends. He also sits on the Hadoop Summit EU data science selection committee, and has spoken at many conferences on a variety of data topics. He also enjoys participating in the Data Science and Big Data communities where he lives in London.Antoine Amend is a data scientist passionate about big data engineering and scalable computing. The books theme of torturing astronomical amounts of unstructured data to gain new insights mainly comes from his background in theoretical physics. Graduating in 2008 with a Msc. in Astrophysics, he worked for a large consultancy business in Switzerland before discovering the concept of big data at the early stages of Hadoop. He has embraced big data technologies ever since, and is now working as the Head of Data Science for cyber security at Barclays Bank. By combining a scientific approach with core IT skills, Antoine qualified two years running for the Big Data World Championships finals held in Austin TX. He Placed in the top 12 in both 2014 and 2015 edition (over 2000+ competitors) where he additionally won the Innovation Award using the methodologies and technologies explained in this book.David George is a distinguished distributed computing expert with 15+ years of data systems experience, mainly with globally recognized IT consultancies and brands. Working with core Hadoop technologies since the early days, he has delivered implementations at the largest scale. David always takes a pragmatic approach to software design and values elegance in simplicity.Today he continues to work as a lead engineer, designing scalable applications for financial sector customers with some of the toughest requirements. His latest projects focus on the adoption of advanced AI techniques for increasing levels of automation across knowledge-based industries.Matthew Hallett is a Software Engineer and Computer Scientist with over 15 years of industry experience. He is an expert Object Oriented programmer and systems engineer with extensive knowledge of low level programming paradigms and, for the last 8 years, has developed an expertise in Hadoop and distributed programming within mission critical environments, comprising multithousandnode data centres. With consultancy experience in distributed algorithms and the implementation of distributed computing architectures, in a variety of languages, Matthew is currently a Consultant Data Engineer in the Data Science & Engineering team at a top four audit firm.Table of ContentsThe Big Data Science EcosystemData AcquisitionInput Formats and SchemaExploratory Data AnalysisSpark for Geographic AnalysisScraping Link-Based External DataBuilding CommunitiesBuilding a Recommendation SystemNews Dictionary and Real-Time Tagging SystemStory De-duplication and MutationAnomaly Detection on Sentiment AnalysisTrendCalculusSecure DataScalable Algorithms

Reviews

Write your review

Recommended