TY - JOUR
T1 - Comparative omics-driven genome annotation refinement
T2 - Application across yersiniae
AU - Schrimpe-Rutledge, Alexandra C.
AU - Jones, Marcus B.
AU - Chauhan, Sadhana
AU - Purvine, Samuel O.
AU - Sanford, James A.
AU - Monroe, Matthew E.
AU - Brewer, Heather M.
AU - Payne, Samuel H.
AU - Ansong, Charles
AU - Frank, Bryan C.
AU - Smith, Richard D.
AU - Peterson, Scott N.
AU - Motin, Vladimir L.
AU - Adkins, Joshua N.
PY - 2012/3/27
Y1 - 2012/3/27
N2 - Genome sequencing continues to be a rapidly evolving technology, yet most downstream aspects of genome annotation pipelines remain relatively stable or are even being abandoned. The annotation process is now performed almost exclusively in an automated fashion to balance the large number of sequences generated. One possible way of reducing errors inherent to automated computational annotations is to apply data from omics measurements (i.e. transcriptional and proteomic) to the un-annotated genome with a proteogenomic-based approach. Here, the concept of annotation refinement has been extended to include a comparative assessment of genomes across closely related species. Transcriptomic and proteomic data derived from highly similar pathogenic Yersiniae (Y. pestis CO92, Y. pestis Pestoides F, and Y. pseudotuberculosis PB1/+) was used to demonstrate a comprehensive comparative omic-based annotation methodology. Peptide and oligo measurements experimentally validated the expression of nearly 40% of each strain's predicted proteome and revealed the identification of 28 novel and 68 incorrect (i.e., observed frameshifts, extended start sites, and translated pseudogenes) protein-coding sequences within the three current genome annotations. Gene loss is presumed to play a major role in Y. pestis acquiring its niche as a virulent pathogen, thus the discovery of many translated pseudogenes, including the insertion-ablated argD, underscores a need for functional analyses to investigate hypotheses related to divergence. Refinements included the discovery of a seemingly essential ribosomal protein, several virulence-associated factors, a transcriptional regulator, and many hypothetical proteins that were missed during annotation.
AB - Genome sequencing continues to be a rapidly evolving technology, yet most downstream aspects of genome annotation pipelines remain relatively stable or are even being abandoned. The annotation process is now performed almost exclusively in an automated fashion to balance the large number of sequences generated. One possible way of reducing errors inherent to automated computational annotations is to apply data from omics measurements (i.e. transcriptional and proteomic) to the un-annotated genome with a proteogenomic-based approach. Here, the concept of annotation refinement has been extended to include a comparative assessment of genomes across closely related species. Transcriptomic and proteomic data derived from highly similar pathogenic Yersiniae (Y. pestis CO92, Y. pestis Pestoides F, and Y. pseudotuberculosis PB1/+) was used to demonstrate a comprehensive comparative omic-based annotation methodology. Peptide and oligo measurements experimentally validated the expression of nearly 40% of each strain's predicted proteome and revealed the identification of 28 novel and 68 incorrect (i.e., observed frameshifts, extended start sites, and translated pseudogenes) protein-coding sequences within the three current genome annotations. Gene loss is presumed to play a major role in Y. pestis acquiring its niche as a virulent pathogen, thus the discovery of many translated pseudogenes, including the insertion-ablated argD, underscores a need for functional analyses to investigate hypotheses related to divergence. Refinements included the discovery of a seemingly essential ribosomal protein, several virulence-associated factors, a transcriptional regulator, and many hypothetical proteins that were missed during annotation.
UR - http://www.scopus.com/inward/record.url?scp=84858988418&partnerID=8YFLogxK
UR - http://www.scopus.com/inward/citedby.url?scp=84858988418&partnerID=8YFLogxK
U2 - 10.1371/journal.pone.0033903
DO - 10.1371/journal.pone.0033903
M3 - Article
C2 - 22479471
AN - SCOPUS:84858988418
SN - 1932-6203
VL - 7
JO - PloS one
JF - PloS one
IS - 3
M1 - e33903
ER -