• français
    • English
  • English 
    • français
    • English
  • Login
JavaScript is disabled for your browser. Some features of this site may not work without it.
BIRD Home

Browse

This CollectionBy Issue DateAuthorsTitlesSubjectsJournals BIRDResearch centres & CollectionsBy Issue DateAuthorsTitlesSubjectsJournals

My Account

Login

Statistics

View Usage Statistics

Deep Learning for Metagenomic Data: using 2D Embeddings and Convolutional Neural Networks

Thumbnail
Date
2017
Link to item file
https://hal.sorbonne-universite.fr/hal-01783588
Dewey
Informatique générale
Sujet
Deep Learning
Conference name
NIPS 2017 Workshop on Machine Learning for Health
Conference date
12-2017
Conference city
Long Beach, CA
Conference country
UNITED STATES
URI
https://basepub.dauphine.fr/handle/123456789/21016
Collections
  • LAMSADE : Publications
Metadata
Show full item record
Author
Thanh Hai, Nguyen
Chevaleyre, Yann
Prifti, Edi
Sokolovska, Nataliya
Zucker, Jean-Daniel
Type
Communication / Conférence
Abstract (EN)
Deep learning (DL) techniques have had unprecedented success when applied to images, waveforms, and texts to cite a few. In general, when the sample size (N) is much greater than the number of features (d), DL outperforms previous machine learning (ML) techniques, often through the use of convolution neural networks (CNNs). However, in many bioinformatics ML tasks, we encounter the opposite situation where d is greater than N. In these situations, applying DL techniques (such as feed-forward networks) would lead to severe overfitting. Thus, sparse ML techniques (such as LASSO e.g.) usually yield the best results on these tasks. In this paper, we show how to apply CNNs on data which do not have originally an image structure (in particular on metagenomic data). Our first contribution is to show how to map metagenomic data in a meaningful way to 1D or 2D images. Based on this representation, we then apply a CNN, with the aim of predicting various diseases. The proposed approach is applied on six different datasets including in total over 1000 samples from various diseases. This approach could be a promising one for prediction tasks in the bioinformatics field.

  • Accueil Bibliothèque
  • Site de l'Université Paris-Dauphine
  • Contact
SCD Paris Dauphine - Place du Maréchal de Lattre de Tassigny 75775 Paris Cedex 16

 Content on this site is licensed under a Creative Commons 2.0 France (CC BY-NC-ND 2.0) license.