Difference between revisions of "BIN-Data integration"
m |
m |
||
Line 1: | Line 1: | ||
<div id="ABC"> | <div id="ABC"> | ||
− | <div style="padding:5px; border:1px solid #000000; background-color:# | + | <div style="padding:5px; border:1px solid #000000; background-color:#b3dbce; font-size:300%; font-weight:400; color: #000000; width:100%;"> |
Data Integration | Data Integration | ||
− | <div style="padding:5px; margin-top:20px; margin-bottom:10px; background-color:# | + | <div style="padding:5px; margin-top:20px; margin-bottom:10px; background-color:#b3dbce; font-size:30%; font-weight:200; color: #000000; "> |
(Integration of biological data; Identifier mapping; Entrez; UniProt; BioMart. ID mapping service and match() function.) | (Integration of biological data; Identifier mapping; Entrez; UniProt; BioMart. ID mapping service and match() function.) | ||
</div> | </div> | ||
Line 10: | Line 10: | ||
− | <div style="padding:5px; border:1px solid #000000; background-color:# | + | <div style="padding:5px; border:1px solid #000000; background-color:#b3dbce33; font-size:85%;"> |
<div style="font-size:118%;"> | <div style="font-size:118%;"> | ||
<b>Abstract:</b><br /> | <b>Abstract:</b><br /> | ||
Line 62: | Line 62: | ||
− | |||
{{Smallvspace}} | {{Smallvspace}} | ||
Line 137: | Line 136: | ||
[[Category:ABC-units]] | [[Category:ABC-units]] | ||
{{UNIT}} | {{UNIT}} | ||
− | {{ | + | {{LIVE}} |
</div> | </div> | ||
<!-- [END] --> | <!-- [END] --> |
Latest revision as of 16:32, 24 September 2020
Data Integration
(Integration of biological data; Identifier mapping; Entrez; UniProt; BioMart. ID mapping service and match() function.)
Abstract:
Data integration is a challenging problem. This unit discusses the issues and how the large databases solve this with NCBI's Entrez system and the EBI's UniProt Knoledeg Base and BioMart System. R coding exercises put some technical issues in practice.
Objectives:
|
Outcomes:
|
Deliverables:
Prerequisites:
This unit builds on material covered in the following prerequisite units:
Evaluation
Evaluation: NA
Contents
Task:
- Read the introductory notes on concepts and approaches to data integration in bioinformatics.
Task:
- Visit the UniProt ID mapping service, enter
NP_010227
into the identifier field, select options from RefSeq Protein to UniProtKB and click Go. - Confirm that this retrieved the right identifier.
- Also note that you could have searched with a list of IDs, and downloaded the results, e.g. for further processing in R.
Task:
- Open RStudio and load the
ABC-units
R project. If you have loaded it before, choose File → Recent projects → ABC-Units. If you have not loaded it before, follow the instructions in the RPR-Introduction unit. - Choose Tools → Version Control → Pull Branches to fetch the most recent version of the project from its GitHub repository with all changes and bug fixes included.
- Type
init()
if requested. - Open the file
BIN-Data_integration.R
and follow the instructions.
Note: take care that you understand all of the code in the script. Evaluation in this course is cumulative and you may be asked to explain any part of code.
Task:
The biomartr
bioconductor package is a second-generation R interface to BioMart that extends the biomaRt
package. It has a good quick start introduction to "Functional Annotation".
- Navigate to https://cran.r-project.org/web/packages/biomartr/vignettes/Functional_Annotation.html
- Work through the tutorial.
Further reading, links and resources
Xie & Ahn (2010) Statistical methods for integrating multiple types of high-throughput data. Methods Mol Biol 620:511-29. (pmid: 20652519) |
Notes
About ...
Author:
- Boris Steipe <boris.steipe@utoronto.ca>
Created:
- 2017-08-05
Modified:
- 2020-09-24
Version:
- 1.1
Version history:
- 1.1 2020 Maintenance
- 1.0 First live version.
- 0.1 First stub
This copyrighted material is licensed under a Creative Commons Attribution 4.0 International License. Follow the link to learn more.