Big Data Processing Using Spark in Cloud -

Big Data Processing Using Spark in Cloud (eBook)

eBook Download: PDF
2018 | 1st ed. 2019
XIII, 264 Seiten
Springer Singapore (Verlag)
978-981-13-0550-4 (ISBN)
Systemvoraussetzungen
96,29 inkl. MwSt
  • Download sofort lieferbar
  • Zahlungsarten anzeigen

The book describes the emergence of big data technologies and the role of Spark in the entire big data stack. It compares Spark and Hadoop and identifies the shortcomings of Hadoop that have been overcome by Spark. The book mainly focuses on the in-depth architecture of Spark and our understanding of Spark RDDs and how RDD complements big data's immutable nature, and solves it with lazy evaluation, cacheable and type inference. It also addresses advanced topics in Spark, starting with the basics of Scala and the core Spark framework, and exploring Spark data frames, machine learning using Mllib, graph analytics using Graph X and real-time processing with Apache Kafka, AWS Kenisis, and Azure Event Hub. It then goes on to investigate Spark using PySpark and R. Focusing on the current big data stack, the book examines the interaction with current big data tools, with Spark being the core processing layer for all types of data.

The book is intended for data engineers and scientists working on massive datasets and big data technologies in the cloud. In addition to industry professionals, it is helpful for aspiring data processing professionals and students working in big data processing and cloud computing environments.



Mamta Mittal, Ph.D., is currently working at G.B. Pant Govt. Engineering College, Okhla, New Delhi. She graduated with a degree in Computer Science & Engineering from Kurukshetra University and received her Master's degree (Honors) in Computer Science & Engineering from YMCA, Faridabad. She subsequently completed her Ph.D. in Computer Science and Engineering at Thapar University, Patiala.  She has been teaching for the past 15 years with a focus on data mining, DBMS, operating systems and data structures. She is an active member of the CSI and IEEE.

Valentina E. Balas, Ph.D., is currently a Full Professor at the Department of Automatics and Applied Software at the Faculty of Engineering, 'Aurel Vlaicu' University of Arad, Romania. She holds a Ph.D. in Applied Electronics and Telecommunications from the Polytechnic University of Timisoara. Dr. Balas is the author of more than 270 research papers in refereed journals and for international conferences. Her research interests are in intelligent systems, fuzzy control, soft computing, smart sensors, information fusion, modeling and simulation. She is the Editor-in-Chief of the International Journal of Advanced Intelligence Paradigms (IJAIP) and International Journal of Computational Systems Engineering (IJCSysE), serves on the Editorial Board of several national and international journals, and as an evaluator expert for national and international projects. She was General Chair of the International Workshop on Soft Computing and Applications held in Romania and Hungary (2005-2016).

Lalit Mohan Goyal, Ph.D., received his B.Tech (Honors) in Computer Science & Engineering from Kurukshetra University, his M.Tech (Honors) in Information Technology from Guru Gobind Singh Indraprastha University, New Delhi, and his Ph.D. in Computer Engineering from Jamia Millia Islamia, New Delhi. He has 14 years of teaching experience in the areas of parallel and random algorithms and theory of computation. Presently, he is working at Bharati Vidyapeeth's College of Engineering, New Delhi.

Raghvendra Kumar, Ph.D., is currently an Assistant Professor at the Department of Computer Science and Engineering, LNCT College, Jabalpur, and at Jodhpur National University, Rajasthan, India. He completed his Bachelor of Technology at SRM University, Chennai and his Master of Technology at KIIT University, Odisha. His research interests include graph theory, discrete mathematics, robotics, cloud computing and algorithms. He also works as a reviewer, and an editorial and technical board member for various journals.


The book describes the emergence of big data technologies and the role of Spark in the entire big data stack. It compares Spark and Hadoop and identifies the shortcomings of Hadoop that have been overcome by Spark. The book mainly focuses on the in-depth architecture of Spark and our understanding of Spark RDDs and how RDD complements big data's immutable nature, and solves it with lazy evaluation, cacheable and type inference. It also addresses advanced topics in Spark, starting with the basics of Scala and the core Spark framework, and exploring Spark data frames, machine learning using Mllib, graph analytics using Graph X and real-time processing with Apache Kafka, AWS Kenisis, and Azure Event Hub. It then goes on to investigate Spark using PySpark and R. Focusing on the current big data stack, the book examines the interaction with current big data tools, with Spark being the core processing layer for all types of data.The book is intended for data engineers and scientistsworking on massive datasets and big data technologies in the cloud. In addition to industry professionals, it is helpful for aspiring data processing professionals and students working in big data processing and cloud computing environments.

Mamta Mittal, Ph.D., is currently working at G.B. Pant Govt. Engineering College, Okhla, New Delhi. She graduated with a degree in Computer Science & Engineering from Kurukshetra University and received her Master’s degree (Honors) in Computer Science & Engineering from YMCA, Faridabad. She subsequently completed her Ph.D. in Computer Science and Engineering at Thapar University, Patiala.  She has been teaching for the past 15 years with a focus on data mining, DBMS, operating systems and data structures. She is an active member of the CSI and IEEE. Valentina E. Balas, Ph.D., is currently a Full Professor at the Department of Automatics and Applied Software at the Faculty of Engineering, “Aurel Vlaicu” University of Arad, Romania. She holds a Ph.D. in Applied Electronics and Telecommunications from the Polytechnic University of Timisoara. Dr. Balas is the author of more than 270 research papers in refereed journals and for international conferences. Her research interests are in intelligent systems, fuzzy control, soft computing, smart sensors, information fusion, modeling and simulation. She is the Editor-in-Chief of the International Journal of Advanced Intelligence Paradigms (IJAIP) and International Journal of Computational Systems Engineering (IJCSysE), serves on the Editorial Board of several national and international journals, and as an evaluator expert for national and international projects. She was General Chair of the International Workshop on Soft Computing and Applications held in Romania and Hungary (2005-2016). Lalit Mohan Goyal, Ph.D., received his B.Tech (Honors) in Computer Science & Engineering from Kurukshetra University, his M.Tech (Honors) in Information Technology from Guru Gobind Singh Indraprastha University, New Delhi, and his Ph.D. in Computer Engineering from Jamia Millia Islamia, New Delhi. He has 14 years of teaching experience in the areas of parallel and random algorithms and theory of computation. Presently, he is working at Bharati Vidyapeeth’s College of Engineering, New Delhi. Raghvendra Kumar, Ph.D., is currently an Assistant Professor at the Department of Computer Science and Engineering, LNCT College, Jabalpur, and at Jodhpur National University, Rajasthan, India. He completed his Bachelor of Technology at SRM University, Chennai and his Master of Technology at KIIT University, Odisha. His research interests include graph theory, discrete mathematics, robotics, cloud computing and algorithms. He also works as a reviewer, and an editorial and technical board member for various journals.

Concepts of Big Data and Apache Spark.- Big Data Analysis in Cloud and Machine Learning.- Security Issues and Challenges related to Big Data.- Big Data Security Solutions in Cloud.- Data Science and Analytics.- Big Data Technologies.- Data Analysis with Casandra and Spark.- Spin up the Spark Cluster.- Learn Scala.- IO for Spark.- Processing with Spark.- Spark Data Frames and Spark SQL.- Machine Learning and Advanced Analytics.- Parallel Programming with Spark.- Distributed Graph Processing with Spark.- Real Time Processing with Spark.- Spark in Real World.- Case Studies. 

Erscheint lt. Verlag 16.6.2018
Reihe/Serie Studies in Big Data
Studies in Big Data
Zusatzinfo XIII, 264 p. 89 illus., 62 illus. in color.
Verlagsort Singapore
Sprache englisch
Themenwelt Informatik Datenbanken Data Warehouse / Data Mining
Informatik Netzwerke Sicherheit / Firewall
Mathematik / Informatik Informatik Software Entwicklung
Mathematik / Informatik Mathematik Finanz- / Wirtschaftsmathematik
Naturwissenschaften
Wirtschaft
Schlagworte Big data analysis • Casandra • Cloud Computing • Data Analysis • Data processing • privacy preservation • Spark Cluster • Spark SQL
ISBN-10 981-13-0550-1 / 9811305501
ISBN-13 978-981-13-0550-4 / 9789811305504
Haben Sie eine Frage zum Produkt?
PDFPDF (Wasserzeichen)
Größe: 8,8 MB

DRM: Digitales Wasserzeichen
Dieses eBook enthält ein digitales Wasser­zeichen und ist damit für Sie persona­lisiert. Bei einer missbräuch­lichen Weiter­gabe des eBooks an Dritte ist eine Rück­ver­folgung an die Quelle möglich.

Dateiformat: PDF (Portable Document Format)
Mit einem festen Seiten­layout eignet sich die PDF besonders für Fach­bücher mit Spalten, Tabellen und Abbild­ungen. Eine PDF kann auf fast allen Geräten ange­zeigt werden, ist aber für kleine Displays (Smart­phone, eReader) nur einge­schränkt geeignet.

Systemvoraussetzungen:
PC/Mac: Mit einem PC oder Mac können Sie dieses eBook lesen. Sie benötigen dafür einen PDF-Viewer - z.B. den Adobe Reader oder Adobe Digital Editions.
eReader: Dieses eBook kann mit (fast) allen eBook-Readern gelesen werden. Mit dem amazon-Kindle ist es aber nicht kompatibel.
Smartphone/Tablet: Egal ob Apple oder Android, dieses eBook können Sie lesen. Sie benötigen dafür einen PDF-Viewer - z.B. die kostenlose Adobe Digital Editions-App.

Buying eBooks from abroad
For tax law reasons we can sell eBooks just within Germany and Switzerland. Regrettably we cannot fulfill eBook-orders from other countries.

Mehr entdecken
aus dem Bereich
Datenschutz und Sicherheit in Daten- und KI-Projekten

von Katharine Jarmul

eBook Download (2024)
O'Reilly Verlag
39,90