Multi-Resolution Weak Supervision for Sequential Data

10/21/2019
by   Frederic Sala, et al.
16

Since manually labeling training data is slow and expensive, recent industrial and scientific research efforts have turned to weaker or noisier forms of supervision sources. However, existing weak supervision approaches fail to model multi-resolution sources for sequential data, like video, that can assign labels to individual elements or collections of elements in a sequence. A key challenge in weak supervision is estimating the unknown accuracies and correlations of these sources without using labeled data. Multi-resolution sources exacerbate this challenge due to complex correlations and sample complexity that scales in the length of the sequence. We propose Dugong, the first framework to model multi-resolution weak supervision sources with complex correlations to assign probabilistic labels to training data. Theoretically, we prove that Dugong, under mild conditions, can uniquely recover the unobserved accuracy and correlation parameters and use parameter sharing to improve sample complexity. Our method assigns clinician-validated labels to population-scale biomedical video repositories, helping outperform traditional supervision by 36.8 F1 points and addressing a key use case where machine learning has been severely limited by the lack of expert labeled data. On average, Dugong improves over traditional supervision by 16.0 F1 points and existing weak supervision approaches by 24.2 F1 points across several video and sensor classification tasks.

READ FULL TEXT

page 5

page 11

page 12

page 14

page 15

page 20

page 21

page 25

research
03/14/2019

Learning Dependency Structures for Weak Supervision Models

Labeling training data is a key bottleneck in the modern machine learnin...
research
10/05/2018

Training Complex Models with Multi-Task Weak Supervision

As machine learning models continue to increase in complexity, collectin...
research
12/02/2018

Snorkel DryBell: A Case Study in Deploying Weak Supervision at Industrial Scale

Labeling training data is one of the most costly bottlenecks in developi...
research
03/24/2022

Shoring Up the Foundations: Fusing Model Embeddings and Weak Supervision

Foundation models offer an exciting new paradigm for constructing models...
research
02/11/2022

A Survey on Programmatic Weak Supervision

Labeling training data has become one of the major roadblocks to using m...
research
09/07/2017

Inferring Generative Model Structure with Static Analysis

Obtaining enough labeled data to robustly train complex discriminative m...
research
03/02/2017

Learning the Structure of Generative Models without Labeled Data

Curating labeled training data has become the primary bottleneck in mach...

Please sign up or login with your details

Forgot password? Click here to reset