Logo image
Generalized Similarity Kernels for Efficient Sequence Classification
Technical documentation   Open access

Generalized Similarity Kernels for Efficient Sequence Classification

Pavel Kuksa, Imdadullah Khan and Vladimir Pavlovic
Rutgers University
2011
DOI:
https://doi.org/10.7282/T3280C58

Abstract

String kernel-based machine learning methods have yielded great success in practical tasks of structured/sequential data analysis. In this paper we propose a novel computational framework that uses general similarity metrics and distance-preserving embeddings with string kernels to improve sequence classification. An embedding step, a distance-preserving bitstring mapping, is used to effectively capture similarity between otherwise symbolically different sequence elements. We show that it is possible to retain computational efficiency of string kernels while using this more “precise” measure of similarity. We then demonstrate that on a number of sequence classification tasks such as music, and biological sequence classification, the new method can substantially improve upon state-of-the-art string kernel baselines.
pdf
tr5b527ef9e1ffe164.24 kBDownloadView
Version of Record (VoR) Technical Documentation Open Access
url
Report an accessibility issueView
Please complete a content remediation request to report an accessibility issue with a library electronic resource, website, or service.

Metrics

186 File downloads
69 Record Views

Details

Logo image