Large-scale machine learning for metagenomics sequence classification
Identifieur interne : 001721 ( Main/Exploration ); précédent : 001720; suivant : 001722Large-scale machine learning for metagenomics sequence classification
Auteurs : Kévin Vervier ; Pierre Mahé ; Maud Tournoud ; Jean-Baptiste Veyrieras ; Jean-Philippe VertSource :
- Bioinformatics [ 1367-4803 ] ; 2015.
Descripteurs français
- KwdFr :
- MESH :
English descriptors
- KwdEn :
- MESH :
Abstract
Url:
DOI: 10.1093/bioinformatics/btv683
PubMed: 26589281
PubMed Central: 4896366
Affiliations:
Links toward previous steps (curation, corpus...)
- to stream Pmc, to step Corpus: 000085
- to stream Pmc, to step Curation: 000085
- to stream Pmc, to step Checkpoint: 000E10
- to stream PubMed, to step Corpus: 001382
- to stream PubMed, to step Curation: 001382
- to stream PubMed, to step Checkpoint: 001031
- to stream Ncbi, to step Merge: 001360
- to stream Ncbi, to step Curation: 001360
- to stream Ncbi, to step Checkpoint: 001360
- to stream Main, to step Merge: 001726
- to stream Main, to step Curation: 001721
Le document en format XML
<record><TEI><teiHeader><fileDesc><titleStmt><title xml:lang="en">Large-scale machine learning for metagenomics sequence classification</title>
<author><name sortKey="Vervier, Kevin" sort="Vervier, Kevin" uniqKey="Vervier K" first="Kévin" last="Vervier">Kévin Vervier</name>
<affiliation><nlm:aff id="btv683-AFF1"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF2"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF3"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF4"></nlm:aff>
</affiliation>
</author>
<author><name sortKey="Mahe, Pierre" sort="Mahe, Pierre" uniqKey="Mahe P" first="Pierre" last="Mahé">Pierre Mahé</name>
<affiliation><nlm:aff id="btv683-AFF1"></nlm:aff>
</affiliation>
</author>
<author><name sortKey="Tournoud, Maud" sort="Tournoud, Maud" uniqKey="Tournoud M" first="Maud" last="Tournoud">Maud Tournoud</name>
<affiliation><nlm:aff id="btv683-AFF1"></nlm:aff>
</affiliation>
</author>
<author><name sortKey="Veyrieras, Jean Baptiste" sort="Veyrieras, Jean Baptiste" uniqKey="Veyrieras J" first="Jean-Baptiste" last="Veyrieras">Jean-Baptiste Veyrieras</name>
<affiliation><nlm:aff id="btv683-AFF1"></nlm:aff>
</affiliation>
</author>
<author><name sortKey="Vert, Jean Philippe" sort="Vert, Jean Philippe" uniqKey="Vert J" first="Jean-Philippe" last="Vert">Jean-Philippe Vert</name>
<affiliation><nlm:aff id="btv683-AFF2"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF3"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF4"></nlm:aff>
</affiliation>
</author>
</titleStmt>
<publicationStmt><idno type="wicri:source">PMC</idno>
<idno type="pmid">26589281</idno>
<idno type="pmc">4896366</idno>
<idno type="url">http://www.ncbi.nlm.nih.gov/pmc/articles/PMC4896366</idno>
<idno type="RBID">PMC:4896366</idno>
<idno type="doi">10.1093/bioinformatics/btv683</idno>
<date when="2015">2015</date>
<idno type="wicri:Area/Pmc/Corpus">000085</idno>
<idno type="wicri:explorRef" wicri:stream="Pmc" wicri:step="Corpus" wicri:corpus="PMC">000085</idno>
<idno type="wicri:Area/Pmc/Curation">000085</idno>
<idno type="wicri:explorRef" wicri:stream="Pmc" wicri:step="Curation">000085</idno>
<idno type="wicri:Area/Pmc/Checkpoint">000E10</idno>
<idno type="wicri:explorRef" wicri:stream="Pmc" wicri:step="Checkpoint">000E10</idno>
<idno type="wicri:source">PubMed</idno>
<idno type="RBID">pubmed:26589281</idno>
<idno type="wicri:Area/PubMed/Corpus">001382</idno>
<idno type="wicri:explorRef" wicri:stream="PubMed" wicri:step="Corpus" wicri:corpus="PubMed">001382</idno>
<idno type="wicri:Area/PubMed/Curation">001382</idno>
<idno type="wicri:explorRef" wicri:stream="PubMed" wicri:step="Curation">001382</idno>
<idno type="wicri:Area/PubMed/Checkpoint">001031</idno>
<idno type="wicri:explorRef" wicri:stream="Checkpoint" wicri:step="PubMed">001031</idno>
<idno type="wicri:Area/Ncbi/Merge">001360</idno>
<idno type="wicri:Area/Ncbi/Curation">001360</idno>
<idno type="wicri:Area/Ncbi/Checkpoint">001360</idno>
<idno type="wicri:doubleKey">1367-4803:2015:Vervier K:large:scale:machine</idno>
<idno type="wicri:Area/Main/Merge">001726</idno>
<idno type="wicri:Area/Main/Curation">001721</idno>
<idno type="wicri:Area/Main/Exploration">001721</idno>
</publicationStmt>
<sourceDesc><biblStruct><analytic><title xml:lang="en" level="a" type="main">Large-scale machine learning for metagenomics sequence classification</title>
<author><name sortKey="Vervier, Kevin" sort="Vervier, Kevin" uniqKey="Vervier K" first="Kévin" last="Vervier">Kévin Vervier</name>
<affiliation><nlm:aff id="btv683-AFF1"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF2"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF3"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF4"></nlm:aff>
</affiliation>
</author>
<author><name sortKey="Mahe, Pierre" sort="Mahe, Pierre" uniqKey="Mahe P" first="Pierre" last="Mahé">Pierre Mahé</name>
<affiliation><nlm:aff id="btv683-AFF1"></nlm:aff>
</affiliation>
</author>
<author><name sortKey="Tournoud, Maud" sort="Tournoud, Maud" uniqKey="Tournoud M" first="Maud" last="Tournoud">Maud Tournoud</name>
<affiliation><nlm:aff id="btv683-AFF1"></nlm:aff>
</affiliation>
</author>
<author><name sortKey="Veyrieras, Jean Baptiste" sort="Veyrieras, Jean Baptiste" uniqKey="Veyrieras J" first="Jean-Baptiste" last="Veyrieras">Jean-Baptiste Veyrieras</name>
<affiliation><nlm:aff id="btv683-AFF1"></nlm:aff>
</affiliation>
</author>
<author><name sortKey="Vert, Jean Philippe" sort="Vert, Jean Philippe" uniqKey="Vert J" first="Jean-Philippe" last="Vert">Jean-Philippe Vert</name>
<affiliation><nlm:aff id="btv683-AFF2"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF3"></nlm:aff>
</affiliation>
<affiliation><nlm:aff id="btv683-AFF4"></nlm:aff>
</affiliation>
</author>
</analytic>
<series><title level="j">Bioinformatics</title>
<idno type="ISSN">1367-4803</idno>
<idno type="eISSN">1367-4811</idno>
<imprint><date when="2015">2015</date>
</imprint>
</series>
</biblStruct>
</sourceDesc>
</fileDesc>
<profileDesc><textClass><keywords scheme="KwdEn" xml:lang="en"><term>Algorithms</term>
<term>Machine Learning</term>
<term>Metagenome</term>
<term>Metagenomics</term>
<term>Sequence Analysis, DNA</term>
<term>Software</term>
</keywords>
<keywords scheme="KwdFr" xml:lang="fr"><term>Algorithmes</term>
<term>Analyse de séquence d'ADN</term>
<term>Apprentissage machine</term>
<term>Logiciel</term>
<term>Métagénome</term>
<term>Métagénomique</term>
</keywords>
<keywords scheme="MESH" xml:lang="en"><term>Algorithms</term>
<term>Machine Learning</term>
<term>Metagenome</term>
<term>Metagenomics</term>
<term>Sequence Analysis, DNA</term>
<term>Software</term>
</keywords>
<keywords scheme="MESH" xml:lang="fr"><term>Algorithmes</term>
<term>Analyse de séquence d'ADN</term>
<term>Apprentissage machine</term>
<term>Logiciel</term>
<term>Métagénome</term>
<term>Métagénomique</term>
</keywords>
</textClass>
</profileDesc>
</teiHeader>
<front><div type="abstract" xml:lang="en"><p><bold>Motivation:</bold>
Metagenomics characterizes the taxonomic diversity of microbial communities by sequencing DNA directly from an environmental sample. One of the main challenges in metagenomics data analysis is the binning step, where each sequenced read is assigned to a taxonomic clade. Because of the large volume of metagenomics datasets, binning methods need fast and accurate algorithms that can operate with reasonable computing requirements. While standard alignment-based methods provide state-of-the-art performance, compositional approaches that assign a taxonomic class to a DNA read based on the <italic>k</italic>
-mers it contains have the potential to provide faster solutions.</p>
<p><bold>Results:</bold>
We propose a new rank-flexible machine learning-based compositional approach for taxonomic assignment of metagenomics reads and show that it benefits from increasing the number of fragments sampled from reference genome to tune its parameters, up to a coverage of about 10, and from increasing the <italic>k</italic>
-mer size to about 12. Tuning the method involves training machine learning models on about 10<sup>8</sup>
samples in 10<sup>7</sup>
dimensions, which is out of reach of standard softwares but can be done efficiently with modern implementations for large-scale machine learning. The resulting method is competitive in terms of accuracy with well-established alignment and composition-based tools for problems involving a small to moderate number of candidate species and for reasonable amounts of sequencing errors. We show, however, that machine learning-based compositional approaches are still limited in their ability to deal with problems involving a greater number of species and more sensitive to sequencing errors. We finally show that the new method outperforms the state-of-the-art in its ability to classify reads from species of lineage absent from the reference database and confirm that compositional approaches achieve faster prediction times, with a gain of 2–17 times with respect to the BWA-MEM short read mapper, depending on the number of candidate species and the level of sequencing noise.</p>
<p><bold>Availability and implementation:</bold>
Data and codes are available at <ext-link ext-link-type="uri" xlink:href="http://cbio.ensmp.fr/largescalemetagenomics">http://cbio.ensmp.fr/largescalemetagenomics</ext-link>
.</p>
<p><bold>Contact:</bold>
<email>pierre.mahe@biomerieux.com</email>
</p>
<p><bold>Supplementary information:</bold>
<ext-link ext-link-type="uri" xlink:href="http://bioinformatics.oxfordjournals.org/lookup/suppl/doi:10.1093/bioinformatics/btv683/-/DC1">Supplementary data</ext-link>
are available at <italic>Bioinformatics</italic>
online.</p>
</div>
</front>
<back><div1 type="bibliography"><listBibl><biblStruct><analytic><author><name sortKey="Agarwal, A" uniqKey="Agarwal A">A. Agarwal</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Angly, F" uniqKey="Angly F">F. Angly</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Balzer, S" uniqKey="Balzer S">S. Balzer</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Beygelzimer, A" uniqKey="Beygelzimer A">A. Beygelzimer</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Bottou, L" uniqKey="Bottou L">L. Bottou</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Bottou, L" uniqKey="Bottou L">L. Bottou</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Gammerman, A" uniqKey="Gammerman A">A. Gammerman</name>
</author>
<author><name sortKey="Vovk, V" uniqKey="Vovk V">V. Vovk</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Hugenholtz, P" uniqKey="Hugenholtz P">P. Hugenholtz</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Huson, D" uniqKey="Huson D">D. Huson</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Korbel, J" uniqKey="Korbel J">J. Korbel</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Koslicki, D" uniqKey="Koslicki D">D. Koslicki</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Langford, J" uniqKey="Langford J">J. Langford</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Li, H" uniqKey="Li H">H. Li</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Li, H" uniqKey="Li H">H. Li</name>
</author>
<author><name sortKey="Durbin, R" uniqKey="Durbin R">R. Durbin</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Lindner, M S" uniqKey="Lindner M">M.S. Lindner</name>
</author>
<author><name sortKey="Renard, B Y" uniqKey="Renard B">B.Y. Renard</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Lukjancenko, O" uniqKey="Lukjancenko O">O. Lukjancenko</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Mande, S" uniqKey="Mande S">S. Mande</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Martin, J" uniqKey="Martin J">J. Martin</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Mchardy, A C" uniqKey="Mchardy A">A.C. McHardy</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Miller, R" uniqKey="Miller R">R. Miller</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Parks, D" uniqKey="Parks D">D. Parks</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Patil, K" uniqKey="Patil K">K. Patil</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Peterson, J" uniqKey="Peterson J">J. Peterson</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Pruitt, K" uniqKey="Pruitt K">K. Pruitt</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Riesenfeld, C" uniqKey="Riesenfeld C">C. Riesenfeld</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Rosen, G" uniqKey="Rosen G">G. Rosen</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Schmieder, R" uniqKey="Schmieder R">R. Schmieder</name>
</author>
<author><name sortKey="Edwards, R" uniqKey="Edwards R">R. Edwards</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Sonnenburg, S" uniqKey="Sonnenburg S">S. Sonnenburg</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Soon, W" uniqKey="Soon W">W. Soon</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Wang, Q" uniqKey="Wang Q">Q. Wang</name>
</author>
</analytic>
</biblStruct>
<biblStruct><analytic><author><name sortKey="Wood, D E" uniqKey="Wood D">D.E. Wood</name>
</author>
<author><name sortKey="Salzberg, S L" uniqKey="Salzberg S">S.L. Salzberg</name>
</author>
</analytic>
</biblStruct>
</listBibl>
</div1>
</back>
</TEI>
<affiliations><list></list>
<tree><noCountry><name sortKey="Mahe, Pierre" sort="Mahe, Pierre" uniqKey="Mahe P" first="Pierre" last="Mahé">Pierre Mahé</name>
<name sortKey="Tournoud, Maud" sort="Tournoud, Maud" uniqKey="Tournoud M" first="Maud" last="Tournoud">Maud Tournoud</name>
<name sortKey="Vert, Jean Philippe" sort="Vert, Jean Philippe" uniqKey="Vert J" first="Jean-Philippe" last="Vert">Jean-Philippe Vert</name>
<name sortKey="Vervier, Kevin" sort="Vervier, Kevin" uniqKey="Vervier K" first="Kévin" last="Vervier">Kévin Vervier</name>
<name sortKey="Veyrieras, Jean Baptiste" sort="Veyrieras, Jean Baptiste" uniqKey="Veyrieras J" first="Jean-Baptiste" last="Veyrieras">Jean-Baptiste Veyrieras</name>
</noCountry>
</tree>
</affiliations>
</record>
Pour manipuler ce document sous Unix (Dilib)
EXPLOR_STEP=$WICRI_ROOT/Sante/explor/MersV1/Data/Main/Exploration
HfdSelect -h $EXPLOR_STEP/biblio.hfd -nk 001721 | SxmlIndent | more
Ou
HfdSelect -h $EXPLOR_AREA/Data/Main/Exploration/biblio.hfd -nk 001721 | SxmlIndent | more
Pour mettre un lien sur cette page dans le réseau Wicri
{{Explor lien |wiki= Sante |area= MersV1 |flux= Main |étape= Exploration |type= RBID |clé= PMC:4896366 |texte= Large-scale machine learning for metagenomics sequence classification }}
Pour générer des pages wiki
HfdIndexSelect -h $EXPLOR_AREA/Data/Main/Exploration/RBID.i -Sk "pubmed:26589281" \ | HfdSelect -Kh $EXPLOR_AREA/Data/Main/Exploration/biblio.hfd \ | NlmPubMed2Wicri -a MersV1
This area was generated with Dilib version V0.6.33. |