Eric H Lyons
Publications
PMID: 17301222;PMCID: PMC1805546;Abstract:
After the most recent tetraploidy in the Arabidopsis lineage, most gene pairs lost one, but not both, of their duplicates. We manually inspected the 3,179 retained gene pairs and their surrounding gene space still present in the genome using a custom-made viewer application. The display of these pairs allowed us to define intragenic conserved noncoding sequences (CNSs), identify exon annotation errors, and discover potentially new genes. Using a strict algorithm to sort high-scoring pair sequences from the bl2seq data, we created a database of 14,944 intragenomic Arabidopsis CNSs. The mean CNS length is 31 bp, ranging from 15 to 285 bp. There are ≈1.7 CNSs associated with a typical gene, and Arabidopsis CNSs are found in all areas around exons, most frequently in the 5′ upstream region. Gene ontology classifications related to transcription, regulation, or "response to..." external or endogenous stimuli, especially hormones, tend to be significantly overrepresented among genes containing a large number of CNSs, whereas protein localization, transport, and metabolism are common among genes with no CNSs. There is a 1.5% overlap between these CNSs and the 218,982 putative RNAs in the Arabidopsis Small RNA Project database, allowing for two mismatches. These CNSs provide a unique set of noncoding sequences enriched for function. CMS function is implied by evolutionary conservation and independently supported because CNS-richness predicts regulatory gene ontology categories. © 2007 by The National Academy of Sciences of the USA.
PMID: 18952863;PMCID: PMC2593677;Abstract:
In addition to the genomes of Arabidopsis (Arabidopsis thaliana) and poplar (Populus trichocarpa), two near-complete rosid genome sequences, grape (Vitis vinifera) and papaya (Carica papaya), have been recently released. The phylogenetic relationship among these four genomes and the placement of their three independent, fractionated tetraploidies sum to a powerful comparative genomic system. CoGe, a platform of multiple whole or near-complete genome sequences, provides an integrative Web-based system to find and align syntenic chromosomal regions and visualize the output in an intuitive and interactive manner. CoGe has been customized to specifically support comparisons among the rosids. Crucial facts and definitions are presented to clearly describe the sorts of biological questions that might be answered in part using CoGe, including patterns of DNA conservation, accuracy of annotation, transposability of individual genes, subfunctionalization and/or fractionation of syntenic gene sets, and conserved noncoding sequence content. This précis of an online tutorial, CoGe with Rosids (http://tinyurl.com/4a23pk), presents sample results graphically. © 2008 American Society of Plant Biologists.