Showing posts with label Databases. Show all posts
Showing posts with label Databases. Show all posts

Rapid GWAS of thousands of phenotypes for 337,000 samples in the UK Biobank — Neale lab

 "The UK Biobank recently released genome-wide association data on ~500,000 individuals. The genotype data for these samples have been cleaned, imputed and released to the scientific community. This public release of data represents an extraordinary advance for genetics, pushing the envelope for data sharing and rapid uptake by the research community. These data will be used for novel discovery of disease-associated genes, in the development of new methods, and to serve as an example for how future efforts in genetics and biology ought to proceed.

To further enhance the value of this resource, we have performed a basic association test on ~337,000 unrelated individuals of British ancestry for over 2,000 of the available phenotypes. We’re making these results available for browsing through several portals, including the Global Biobank Engine where they will appear soon. They are also available for download here.

We have decided not to write a scientific article for publication based on these analyses. Rather, we have described the data processing in a detailed blog post linked to the underlying code repositories. The decision to eschew scientific publication for the basic association analysis is rooted in our view that we will continue to work on and analyze these data and, as a result, writing a paper would not reflect the current state of the scientific work we are performing. Our goal here is to make these results available as quickly as possible, for any geneticist, biologist or curious citizen to explore. This is not to suggest that we will not write any papers on these data, but rather only write papers for those activities that involve novel method development or more complex analytic approaches. A univariate genome-wide association analysis is now a relatively well-established activity, and while the scale of this is a bit grander than before, that in and of itself is a relatively perfunctory activity. Simply put, let the data be free.

We do view these results as likely to change as we continue to refine the quality control analyses and as we continue to dig into the results themselves. Nevertheless, we’ve started to use them in a variety of downstream analyses and for other scientific projects and hope that others find them useful too."



'via Blog this'

Rat Brain-Connectome

 ChemNetDB is an open access constantly evolving neurochemical connectivity database obtained from the rat brain. By utilizing advanced neuroinformatics methods, we have integrated over 50 years of neuroanatomical and neurochemical research on rat brain into a consistent, brain-wide, multi-scale neurochemical cerebral connectome. This connectome differs from the previously suggested neuronal networks of rat brain in two major aspects:

it integrates information on transmitter systems and receptors on the network topology;
it is unbiased and hypothesis-free as the connectome was obtained by consistent and systematic procedures without a priori assumptions."



World's largest autism genome database shines new light on many 'autisms'

Through its research platform on the Google Cloud, Autism Speaks is making all of MSSNG's fully sequenced genomes directly available to researchers free of charge, along with analytic tools. In the coming weeks, the MSSNG team will be uploading an additional 2,000 fully sequenced autism genomes, bringing the total over 7,000.
Currently, more than 90 investigators at 40 academic and medical institutions are using the MSSNG database to advance autism research around the world."



Whole genome sequencing resource identifies 18 new candidate genes for autism spectrum disorder, Nature Neuroscience (2017). DOI: 10.1038/nn.4524 

Largest resource of human protein-protein interactions can help interpret genomic data

Human Interactome network visualized by Cytosc...
Human Interactome network visualized by Cytoscape 2.5. (Photo credit: Wikipedia 

An international research team has developed the largest database of protein-to-protein interaction networks, a resource that can illuminate how numerous disease-associated genes contribute to disease development and progression. Led by investigators at Massachusetts General Hospital (MGH) and the Broad Institute of MIT and Harvard, the team's report on its development of the network called InWeb_InBioMap (InWeb_IM) is receiving advance online publication in Nature Methods.  

A scored human protein–protein interaction network to catalyze genomic interpretation

The database is at Intomics


The harmonizome: a collection of processed datasets gathered to serve and mine knowledge about genes and proteins. - PubMed

Genomics, epigenomics, transcriptomics, proteomics and metabolomics efforts rapidly generate a plethora of data on the activity and levels of biomolecules within mammalian cells. At the same time, curation projects that organize knowledge from the biomedical literature into online databases are expanding. Hence, there is a wealth of information about genes, proteins and their associations, with an urgent need for data integration to achieve better knowledge extraction and data reuse. For this purpose, we developed the Harmonizome: a collection of processed datasets gathered to serve and mine knowledge about genes and proteins from over 70 major online resources. We extracted, abstracted and organized data into ∼72 million functional associations between genes/proteins and their attributes. Such attributes could be physical relationships with other biomolecules, expression in cell lines and tissues, genetic associations with knockout mouse or human phenotypes, or changes in expression after drug treatment. We stored these associations in a relational database along with rich metadata for the genes/proteins, their attributes and the original resources. The freely available Harmonizome web portal provides a graphical user interface, a web service and a mobile app for querying, browsing and downloading all of the collected data. To demonstrate the utility of the Harmonizome, we computed and visualized gene-gene and attribute-attribute similarity networks, and through unsupervised clustering, identified many unexpected relationships by combining pairs of datasets such as the association between kinase perturbations and disease signatures. We also applied supervised machine learning methods to predict novel substrates for kinases, endogenous ligands for G-protein coupled receptors, mouse phenotypes for knockout genes, and classified unannotated transmembrane proteins for likelihood of being ion channels. The Harmonizome is a comprehensive resource of knowledge about genes and proteins, and as such, it enables researchers to discover novel relationships between biological entities, as well as form novel data-driven hypotheses for experimental validation.Database URL:

SZGR2: Schizophrenia Gene Resource

SZGR2: "Schizophrenia Gene Resource 2 (SZGR 2.0) provides a comprehensive online resource for schizophrenia studies. Built in 2009, SZGR curates data from genetics studies, transcriptome studies, and epigenetics studies. Overtime, SZGR has grown from a pure genetics resource to a comprehensive, multi-omics online data house -- a one-stop shop for schizophrenia. Our long term goal is to provide resources and tools for researchers to understand schizophrenia. Our ultimate goal is to improve the life quality of schizophrenia patients and to contribute to successful therapies for patients. Please email us if there are any feedbacks. Click here to browse all data sets. Documentation can be found here."

EHFPI: a database and analysis resource of essential host factors for pathogenic infection.

High-throughput screening and computational technology has greatly
changed the face of microbiology in better understanding pathogen-host
interactions. Genome-wide RNA interference (RNAi) screens have given
rise to a new class of host genes designated as Essential Host Factors
(EHFs), whose knockdown effects significantly influence pathogenic
infections. Therefore, we present the first release of a
manually-curated bioinformatics database and analysis resource EHFPI
(Essential Host Factors for Pathogenic Infection,). EHFPI captures detailed article,
screen, pathogen and phenotype annotation information for a total of
4634 EHF genes of 25 clinically important pathogenic species. Notably,
EHFPI also provides six powerful and data-integrative analysis tools,
i.e. EHF Overlap Analysis, EHF-pathogen Network Analysis, Gene
Enrichment Analysis, Pathogen Interacting Proteins (PIPs) Analysis, Drug
Target Analysis and GWAS Candidate Gene Analysis, which advance the
comprehensive understanding of the biological roles of EHF genes, as in
diverse perspectives of protein-protein interaction network, drug
targets and diseases/traits. The EHFPI web interface provides
appropriate tools that allow efficient query of EHF data and
visualization of custom-made analysis results. EHFPI data and tools
shall keep available without charge and serve the microbiology,
biomedicine and pharmaceutics research communities, to finally
facilitate the development of diagnostics, prophylactics and
therapeutics for human pathogens.

EpimiR: a database of curated mutual regulation between miRNAs and epigenetic modifications.

As two kinds of important transcription regulators, both epigenetic
modification and miRNA can regulate gene expression in a wide range of
complex diseases. Recently, many studies have demonstrated that
epigenetics and miRNA can regulate each other in many biological
processes. For instance, methylation of promoter-associated CpG
dinucleotides (especially in CpG islands) usually correlates with
reduced transcription levels of corresponding miRNAs and subsequently
induces the expression of miRNA target genes. Additionally, histone
modifications have been discovered to play positive or negative roles in
controlling miRNA expression in various normal cells and diseases. On
the other hand, miRNA exerts its curative effects on regulating DNA
methylation or histone modification through directly targeting
epigenetic enzymes or functional protein complexes. Thus, we developed a
comprehensive database named EpimiR to store the experimentally
validated mutual regulations between epigenetic modifications and
miRNAs. The EpimiR database have obtained 1945 regulatory relationships
between 18 types of epigenetic modification (including DNA methylation,
histone acetylation, H3K4me3 and H3K27me3, etc.) and 615 miRNAs across 6
species (including human, mouse, chicken, virus, canine, and
arabidopsis) from nearly 2000 literatures. The records that were stored
in the EpimiR database can be divided into 2 parts: Epi2miR and miR2Epi.
Users can search, submit, and download with a user-friendly interface.

Home :: BrainSpan: Atlas of the Developing Human Brain


The BrainSpan atlas includes the following

  • Developmental Transcriptome:
    RNA sequencing and exon microarray data
    profiling up to sixteen cortical and subcortical structures across the full course of human brain development.

  • Prenatal LMD Microarray:
    High-resolution neuroanatomical transcriptional profiles
    of ~300 distinct structures spanning the entire brain for four
    midgestional prenatal specimens.

  • ISH:
    High-resolution in situ hybridization image data covering selected genes and brain regions
    in developing and adult human brain.

  • Reference Atlas:
    Full color, high-resolution anatomic reference atlases of prenatal and adult human brain.

NIH Launches First Phase of Microbiome Cloud Project

New Public Resource to Help Researchers Explore Data from the NIH Human Microbiome Project
The National Institutes of Health (NIH) has launched the first phase of the Microbiome Cloud Project (MCP), a collaboration with Amazon Web Services that aims to improve access to and analysis of data from the Human Microbiome Project (HMP). Five terabytes of genetic information on the microbes that naturally colonize our bodies—enough information to fill more than 1,000 standard DVDs—are now available as a free public dataset  allowing users to access and analyze the data online. This cloud, or internet-based, storage facilitates analysis by reducing the need for time-consuming downloads.

The microbiome project cloud

BBC News - Public Health England to launch largest cancer database

The world's largest database of cancer patients is being set up in England in an attempt to revolutionise care, Public Health England has announced.
It will collate all the available data on each of the 350,000 new tumours detected in the country each year.
The aim is to use the register to help usher in an era of "personalised medicine" that will see treatments matched to the exact type of cancer a patient has.

EnrichNet: network-based gene set enrichment analysis.


Assessing functional associations between an experimentally derived gene or protein set of interest and a database of known gene/protein sets is a common task in the analysis of large-scale functional genomics data. For this purpose, a frequently used approach is to apply an over-representation-based enrichment analysis. However, this approach has four drawbacks: (i) it can only score functional associations of overlapping gene/proteins sets; (ii) it disregards genes with missing annotations; (iii) it does not take into account the network structure of physical interactions between the gene/protein sets of interest and (iv) tissue-specific gene/protein set associations cannot be recognized.

RESULTS:

To address these limitations, we introduce an integrative analysis approach and web-application called EnrichNet. It combines a novel graph-based statistic with an interactive sub-network visualization to accomplish two complementary goals: improving the prioritization of putative functional gene/protein set associations by exploiting information from molecular interaction networks and tissue-specific gene expression data and enabling a direct biological interpretation of the results. By using the approach to analyse sets of genes with known involvement in human diseases, new pathway associations are identified, reflecting a dense sub-network of interactions between their corresponding proteins.

AVAILABILITY:

EnrichNet is freely available at http://www.enrichnet.org

FindZebra - The search engine for difficult medical cases

 "There are close to 7,000 rare diseases recognized by rare disease organizations. We index over 31,000 documents covering rare and genetic diseases from 10 reputable sources. Given the number of rare diseases and rate of publication, we think FindZebra is a good companion for medical professionals."


Genome-wide atlas of gene enhancers in the brain online

"Future research into the underlying causes of neurological disorders such as autism, epilepsy and schizophrenia, should greatly benefit from a first-of-its-kind atlas of gene-enhancers in the cerebrum (telencephalon). This new atlas, developed by a team led by researchers with the U.S. Department of Energy (DOE)'s Lawrence Berkeley National Laboratory (Berkeley Lab) is a publicly accessible Web-based collection of data that identifies and locates thousands of gene-regulating elements in a region of the brain that is of critical importance for cognition, motor functions and emotion."

The website is here: The VISTA Enhancer Browser is a central resource for experimentally validated human and mouse noncoding fragments with gene enhancer activity as assessed in transgenic mice. 

Case History Database | From Biomed Central

Documenting a patient's case history to inform physicians how the patient has been evaluated and the subsequent progression of his or her disease is arguably the oldest method of communicating medical evidence. And in the 21st century case reports play an equally important role.

Since the launch of Journal of Medical Case Reports in 2007 and the more recent introduction of case reports to the broad-scope journal BMC Research Notes, BioMed Central has acknowledged the value of case reports to the scientific record.  To strengthen this commitment we have developed a valuable new resource – Cases Database, a continuously-updated, freely-accessible database of thousands of medical case reports from multiple publishers, including Springer, BMJ and PubMed Central.
By aggregating case reports and facilitating comparison, Cases Database provides clinicians, researchers, regulators and patients a simple resource to explore content, and identify emerging trends.
Enhanced by Zemanta

Transcriptomine: Nuclear receptor Transcription targets

Transcriptomine is a tool for mining tissue-specific nuclear receptor transcriptomes based on annotated published genome wide transcriptional profiling experiments in the field of nuclear receptor signaling.

Interactome3D from the University of Barcelona

Interactome3D is a web service for the structural annotation of protein-protein interaction networks. Submit your interactions and the server will find all the available structural data for both the single interactors and the interactions themselves. Additionally you can also visualize and download structural information for interactions involving a set of proteins or interactomes for one of the precalculated organisms.

PrePPI: structure-based prediction of protein-protein interactions

PrePPI is a database of predicted and experimentally determined protein-protein interactions (PPIs) for yeast and human. Predicted interactions are assigned a likelihood using a Bayesian framework that combines structural, functional, evolutionary and expression information. The database contains ~2 million predictions including 31,402 for yeast and 317,813 for human that are considered high confidence based on our analysis. Experimentally determined interactions are compiled from a set of public databases (e.g., DIP, IntAct, HPRD, etc). The manuscript "Structure-based prediction of protein-protein interactions on a genome-wide scale" that describes the complete method will be available soon.

Allen Brain Atlas - The human brain whole genome transcriptome


An "all genes, all structures" gene expression survey in multiple adult control brains.
  • > 62,000 gene probes per profile
  • ~ 500 samples per hemisphere across cerebrum, cerebellum and brainstem
  • Data mapped with histology into unified 3-D anatomic framework based on MRI

SORTALLER:Predicting allergens based on Allergen Family

SORTALLER is an online allergen classifier based on allergen family featured peptide (AFFP) dataset and normalized BLAST E-values, which establish the featured vectors for support vector machine (SVM). AFFPs are allergen-specific peptides panned from irredundant allergens and harbor perfect information with noise fragments eliminated because of their similarity to non-allergens. SORTALLER performed significantly better than other existing software and reached a perfect balance with high specificity (98.4%) and sensitivity (98.6%) for discriminating allergenic proteins from several independent datasets of protein sequences of diverse sources, also highlighting with the Matthews correlation coefficient (MCC) as high as 0.970, fast running speed and rapidly predicting a batch of amino acid sequences with a single click.
Enhanced by Zemanta