Highly Parallel Programming with Apache Spark

Tutorials – Apache Spark

Article from Issue 202/2017

Author(s): Ben Everard

Churn through lots of data with cluster computing on Apache's Spark platform.

As a society, we're creating more data than ever before. We're monitoring everything from the planet's weather to the performance of our computers, and we're storing all this information. But how do you process all this data? On a single machine, you can get a few terabytes of disk space and a few hundred gigabytes of memory (at least, you can if your pockets are deep enough), but how do you churn through a petabyte of raw ones and zeros? Basically, you're going to need more than one computer, and you're going to look for a method of running your programs on many machines at the same time: Apache Spark [1].

Before you run off and buy a rack of servers, slow down! We're going to start by introducing Spark on a single machine. Once you've mastered the basics, you can scale up.

Spark is a data processing engine that is often used with Hadoop for managing large amounts of data in a highly distributed manner. If you move forward with Spark, you're probably going to end up with a complete Hadoop setup; however, that's also getting ahead of ourselves. We can start Spark as a standalone service on a single computer.

[...]

Use Express-Checkout link below to read the full article (PDF).

Buy this article as PDF

Download Article PDF now with Express Checkout

Price $2.95
(incl. VAT)

Buy Linux Magazine

SINGLE ISSUES

Print Issues

Digital Issues

SUBSCRIPTIONS

Print Subscriptions

Digital Subscriptions

Support Our Work

Linux Magazine content is made possible with support from readers like you. Please consider contributing when you’ve found an article to be beneficial.

News

Alpine Linux 3.24 Features Fresh Desktops and a Newer Kernel

Alpine Linux , Gnome , Plasma , Security

If you're a fan of Alpine Linux, it's time to upgrade because the latest version has been released with KDE Plasma 6.6, Gnome 50, and Linux kernel 6.18 LTS.
EU Open Source Strategy Plays Key Role in Tech Sovereignty Package

EU , government , open source

Comprehensive measures adopted by the European Commission aim to reduce dependency on non-EU countries.
Linux Foundation Report Indicates AI Driving Tech Hiring

Artificial Inte... , privacy , Security

Within growing security and skills gaps, AI has been found to be a positive driving force behind tech hiring trends in Europe.
United Nations Open Source Portal Goes Live

open source , projects

A new open source portal seeks to coordinate and scale open source efforts across the United Nations system.
KDE Linux Drops AUR

applications , KDE Linux , Security

KDE Linux developers have dropped the Arch User Repository from the build pipeline due to security concerns; other distributions should consider doing the same.
California May Exempt Linux from Its Age-Verification Law

Linux , privacy , SteamOS

After backlash from the Linux community, California may be backing off on its promise to force all operating systems to verify age, but one platform may still have to comply.
Another Logic Bug Found in Linux Kernel

DEBIAN , Fedora , Kernel , Ubuntu , vulnerability

Qualys has discovered a vulnerability in the Linux kernel that can be used to elevate standard user privileges.
Ubuntu Core 26 Offers Game-Changing Enterprise Features

Enterprise Linux , Security , Ubuntu

Ubuntu Core 26 could be a game-changer for organizations looking for increased security and reliability.
AI Flooding the Linux Kernel Security Mailing List

Artificial Inte... , Kernel , Security

AI is giving Linus Torvalds a headache, but not in the way you might think.
Top Priorities for Open Source Pros Seeking a New Job

FOSS , open source

Professional fulfillment tops the list, according to LPI report.

Highly Parallel Programming with Apache Spark

Tutorials – Apache Spark

Buy this article as PDF

Buy Linux Magazine

Related content

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters

Support Our Work

News

Alpine Linux 3.24 Features Fresh Desktops and a Newer Kernel

EU Open Source Strategy Plays Key Role in Tech Sovereignty Package

Linux Foundation Report Indicates AI Driving Tech Hiring

United Nations Open Source Portal Goes Live

KDE Linux Drops AUR

California May Exempt Linux from Its Age-Verification Law

Another Logic Bug Found in Linux Kernel

Ubuntu Core 26 Offers Game-Changing Enterprise Features

AI Flooding the Linux Kernel Security Mailing List

Top Priorities for Open Source Pros Seeking a New Job

Highly Parallel Programming with Apache Spark

Tutorials – Apache Spark

Buy this article as PDF

Buy Linux Magazine

Related content

Subscribe to our Linux Newsletters Find Linux and Open Source Jobs Subscribe to our ADMIN Newsletters

Support Our Work

News

Tag Cloud

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters