Towards the Unseen: Iterative Text Recognition by Distilling from Errors

07/26/2021
by   Ayan Kumar Bhunia, et al.
0

Visual text recognition is undoubtedly one of the most extensively researched topics in computer vision. Great progress have been made to date, with the latest models starting to focus on the more practical "in-the-wild" setting. However, a salient problem still hinders practical deployment – prior arts mostly struggle with recognising unseen (or rarely seen) character sequences. In this paper, we put forward a novel framework to specifically tackle this "unseen" problem. Our framework is iterative in nature, in that it utilises predicted knowledge of character sequences from a previous iteration, to augment the main network in improving the next prediction. Key to our success is a unique cross-modal variational autoencoder to act as a feedback module, which is trained with the presence of textual error distribution data. This module importantly translate a discrete predicted character space, to a continuous affine transformation parameter space used to condition the visual feature map at next iteration. Experiments on common datasets have shown competitive performance over state-of-the-arts under the conventional setting. Most importantly, under the new disjoint setup where train-test labels are mutually exclusive, ours offers the best performance thus showcasing the capability of generalising onto unseen words.

READ FULL TEXT

page 1

page 2

page 3

page 4

research
05/03/2022

Cross-modal Representation Learning for Zero-shot Action Recognition

We present a cross-modal Transformer-based framework, which jointly enco...
research
11/13/2020

Transductive Zero-Shot Learning using Cross-Modal CycleGAN

In Computer Vision, Zero-Shot Learning (ZSL) aims at classifying unseen ...
research
03/30/2022

An Iterative Co-Training Transductive Framework for Zero Shot Learning

In zero-shot learning (ZSL) community, it is generally recognized that t...
research
03/29/2021

StyleMeUp: Towards Style-Agnostic Sketch-Based Image Retrieval

Sketch-based image retrieval (SBIR) is a cross-modal matching problem wh...
research
11/30/2021

Multi-modal Text Recognition Networks: Interactive Enhancements between Visual and Semantic Features

Linguistic knowledge has brought great benefits to scene text recognitio...
research
06/26/2023

TCEIP: Text Condition Embedded Regression Network for Dental Implant Position Prediction

When deep neural network has been proposed to assist the dentist in desi...
research
05/09/2023

Traffic Forecasting on New Roads Unseen in the Training Data Using Spatial Contrastive Pre-Training

New roads are being constructed all the time. However, the capabilities ...

Please sign up or login with your details

Forgot password? Click here to reset