The MADlib Analytics Library or MAD Skills, the SQL

08/21/2012
∙
by   Joe Hellerstein, et al.
∙
0
∙

MADlib is a free, open source library of in-database analytic methods. It provides an evolving suite of SQL-based algorithms for machine learning, data mining and statistics that run at scale within a database engine, with no need for data import/export to other tools. The goal is for MADlib to eventually serve a role for scalable database systems that is similar to the CRAN library for R: a community repository of statistical methods, this time written with scale and parallelism in mind. In this paper we introduce the MADlib project, including the background that led to its beginnings, and the motivation for its open source nature. We provide an overview of the library's architecture and design patterns, and provide a description of various statistical methods in that context. We include performance and speedup results of a core design pattern from one of those methods over the Greenplum parallel DBMS on a modest-sized test cluster. We then report on two initial efforts at incorporating academic research into MADlib, which is one of the project's goals. MADlib is freely available at http://madlib.net, and the project is open for contributions of both new methods, and ports to additional database platforms.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
∙ 01/15/2019

Integrazione di Apache Hive con Spark

English. This document describes the solutions adopted, which arose from...
research
∙ 02/10/2019

ELKI: A large open-source library for data analysis - ELKI Release 0.7.5 "Heidelberg"

This paper documents the release of the ELKI data mining framework, vers...
research
∙ 08/17/2017

Designing and building the mlpack open-source machine learning library

mlpack is an open-source C++ machine learning library with an emphasis o...
research
∙ 12/12/2021

Graph Pattern Matching in GQL and SQL/PGQ

As graph databases become widespread, JTC1 – the committee in joint char...
research
∙ 08/12/2019

Douglas-Quaid -- Open Source Image Matching Library

Security analysts need to classify, search and correlate numerous images...
research
∙ 04/04/2019

Learning Analytics Made in France: The METALproject

This paper presents the METAL project, an ongoing French open Learning A...
research
∙ 09/26/2019

The Stroke Correspondence Problem, Revisited

We revisit the stroke correspondence problem [13,14]. We optimize this a...

Please sign up or login with your details

Forgot password? Click here to reset