Auto-Tag: Tagging-Data-By-Example in Data Lakes

12/11/2021
by   Yeye He, et al.
43

As data lakes become increasingly popular in large enterprises today, there is a growing need to tag or classify data assets (e.g., files and databases) in data lakes with additional metadata (e.g., semantic column-types), as the inferred metadata can enable a range of downstream applications like data governance (e.g., GDPR compliance), and dataset search. Given the sheer size of today's enterprise data lakes with petabytes of data and millions of data assets, it is imperative that data assets can be “auto-tagged”, using lightweight inference algorithms and minimal user input. In this work, we develop Auto-Tag, a corpus-driven approach that automates data-tagging of custom data types in enterprise data lakes. Using Auto-Tag, users only need to provide one example column to demonstrate the desired data-type to tag. Leveraging an index structure built offline using a lightweight scan of the data lake, which is analogous to pre-training in machine learning, Auto-Tag can infer suitable data patterns to best “describe” the underlying “domain” of the given column at an interactive speed, which can then be used to tag additional data of the same “type” in data lakes. The Auto-Tag approach can adapt to custom data-types, and is shown to be both accurate and efficient. Part of Auto-Tag ships as a “custom-classification” feature in a cloud-based data governance and catalog solution Azure Purview.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
12/07/2022

Tag Embedding and Well-defined Intermediate Representation improve Auto-Formulation of Problem Description

In this report, I address auto-formulation of problem description, the t...
research
08/16/2018

2DR: Towards Fine-Grained 2-D RFID Touch Sensing

In this paper, we introduce 2DR, a single RFID tag which can seamlessly ...
research
04/10/2021

Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes

Complex data pipelines are increasingly common in diverse applications s...
research
12/06/2016

Tag Prediction at Flickr: a View from the Darkroom

Automated photo tagging has established itself as one of the most compel...
research
10/10/2020

Tag Recommendation for Online Q A Communities based on BERT Pre-Training Technique

Online Q A and open source communities use tags and keywords to index,...
research
06/09/2021

Auto-tagging of Short Conversational Sentences using Natural Language Processing Methods

In this study, we aim to find a method to auto-tag sentences specific to...
research
07/18/2017

AirCode: Unobtrusive Physical Tags for Digital Fabrication

We present AirCode, a technique that allows the user to tag physically f...

Please sign up or login with your details

Forgot password? Click here to reset