Expression normalization method: Difference between revisions
No edit summary |
No edit summary |
||
| Line 4: | Line 4: | ||
== Common dispersion survey (at reliable TSS cluster level) == |
== Common dispersion survey (at reliable TSS cluster level) == |
||
Ultimately we are going to use TSS expressions, and normalization methods have to be examined at the cluster level. RLE (relative log expression) method proposed by DESeq package has been implemented recently in edgeR in addition to TMM. Here we tested (i) no scaling (ii) RLE, (iii) TMM on the cluster expression. |
Ultimately we are going to use TSS expressions beyond gene level, and normalization methods have to be examined at the cluster level. RLE (relative log expression) method proposed by DESeq package has been implemented recently in edgeR in addition to TMM. Here we tested (i) no scaling (ii) RLE, (iii) TMM on the cluster expression. |
||
=== Method === |
=== Method === |
||
Revision as of 11:54, 29 September 2011
There are several ways on how to normalize CAGE expression data. 'default' way of normalization has been TPM, but we could think of recent proposal such as TMM (implemented in edgeR, genomebiology.com/2010/11/3/R25). Another point to be considered is if we should include chromosome M in the library sizes or not.
Common dispersion survey (at reliable TSS cluster level)
Ultimately we are going to use TSS expressions beyond gene level, and normalization methods have to be examined at the cluster level. RLE (relative log expression) method proposed by DESeq package has been implemented recently in edgeR in addition to TMM. Here we tested (i) no scaling (ii) RLE, (iii) TMM on the cluster expression.
Method
- data - Promotome v1.0 - RC1 (release candidate 1, Sep 26th 2011), from https://fantom5-collaboration.gsc.riken.jp/webdav/home/kawaji/110804-dpi-clusters-expression/ only clusters with max_counts >= 50 were used
- replicates - determined by recognizing "donor","donation","pool","biol_rep","tech_rep", in the sample names.
Results
Common dispersion:
- no scaling of sample sizes: 0.2388622
- RLE: 0.2495027
- TMM: 0.2749048
Interpretation:
- no scaling is the best.
- RLE is reasonably good, however no scaling is the best in terms of common dispersion. Only if RLE is proven to rescue some 'poor quality' libraries, RLE can be a reasonable choice.
- TMM is the worst, but it might be improved by taking a better reference samples (at this moment, the first sample is selected for reference). However, the improvement would not be so drastic from our experiences below.
Note:
- quantile-based scaling implemented in edgeR crashed with the data.
Common dispersion survey (at gene level)
Marina performed a series of experiments on TMM normalization (edgeR, genomebiology.com/2010/11/3/R25) with taking the shallowest, the deepest, or universal RNA as a reference and with/without chromosome M. And she simply checked common dispersion on those combinations.
As a result, she found that the best performance (smallest common dispersion) is:
- conventional TPM is recommended (not TMM normalization). Removal of chromosome M from the library size would be a good option.
Here is the methods and the result by Marina Lizio (m.lizio@gsc.riken.jp) August 1, 2011
A-Dataset
I selected a clean subset of human primary cells with 3 replicates (CD14 set, replicates as follows)
Monocytes monocytes mock treated monocytes treated with B glucan monocytes reated with BCG monocytes treated with Candida monocytes treated with Cryptococcus monocytes treated with Group A streptococci monocytes treated with IFN N hexane monocytes treated with Salmonella monocytes treated with Trehalose dimycolate (TDM) monocytes treated with lipopolysaccharide CD14-CD16+ Monocytes CD14+CD16- Monocytes Universal RNA (Clontech) HeLa control for Automation1
As a first trial I chose the RefSeq gene based expression calculated on the raw counts.
B-Parameters tested
Application of TMM normalization with respect to a reference column used for the normalization itself:
- deepest sample **shallowest sample **universal RNA sample **control sample (HeLa for automation)
And exclusion of chrM mapped tags form the dataset prior to normalization
C-Method description
Use of edgeR bioconductor package (version 2.2.5, bioconductor version 2,8, R version 2.13).
- Steps 1-read the gene expression table, read the library sizes
- 2-select the CD14 subset of replicates
- 3-include one control and universal RNA samples for TMM normalization.
- Step3 is done only when TMM normalization is taken into account.
- 4-calculate the normalization factor for all 4 cases (shallowest sample, deepest sample, universal RNA sample, HeLa control sample)
- Step 4 is skipped in the test case without TMM normalization applied.
- 5-calculate the common dispersion of the resulting normalized dataset.
These steps above are repeated for the library sizes without considering chrM tags
The normalization method assigns a normalization factor to each sample, according to the reference sample used. For all the tests the common dispersion was calculated, to check whether a best combination of parameters as in B) exists so that the dispersion is minimized.
[Pseudocode]
inputs=DGEList(infiles, groups=my.groups, libsizes=my.counts)
#the groups are listed in A) and the counts are either the size of the whole
#library or the size of the library without the number of chrM associated tags
#case of TMM normalization
for dex in inputs { for (mycol in c(idx_shallowest,idx_ctrl, idx_deepest,idx_univ)) {
d=calcNormFactors(dex,method="TMM",refColumn=mycol)
d.comm=estimateCommonDisp(d)
} }
#case of no TMM normalization
for dex in inputs { for (mycol in c(idx_shallowest,idx_ctrl, idx_deepest,idx_univ)) {
d.comm=estimateCommonDisp(d)
} }
D-Results
In all cases of normalization, the dispersion is minimized if the universal RNA sample is considered as reference. Dispersion is further improved if the chrM counts are not considered in the library size. However, if no TMM normalization is performed, dispersion is minimized.
#with TMM
> load("CD14_dispersion_replicates_no-chrM_with_control_refcol-hela_control-20110726.RData")
> sqrt(d.comm$common.dispersion)
[1] 0.2419791
> load("CD14_dispersion_replicates_with_control_refcol-hela_control-20110726.RData")
> sqrt(d.comm$common.dispersion)
[1] 0.243903
> load("CD14_dispersion_replicates_with_control_refcol-shallowest_lib-20110726.RData")
> sqrt(d.comm$common.dispersion)
[1] 0.2351437
> load("CD14_dispersion_replicates_no-chrM_with_control_refcol-shallowest_lib-20110726.RData")
> sqrt(d.comm$common.dispersion)
[1] 0.233168
> load("CD14_dispersion_replicates_no-chrM_with_control_refcol-deepest_lib-20110726.RData")
> sqrt(d.comm$common.dispersion)
[1] 0.2397358
> load("CD14_dispersion_replicates_with_control_refcol-deepest_lib-20110726.RData")
> sqrt(d.comm$common.dispersion)
[1] 0.241812
> load("CD14_dispersion_replicates_with_control_refcol-universalRNA-20110726.RData")
> sqrt(d.comm$common.dispersion)
[1] 0.2340208
> load("CD14_dispersion_replicates_no-chrM_with_control_refcol-universalRNA-20110726.RData")
> sqrt(d.comm$common.dispersion)
[1] 0.2321581
#Without TMM
> sqrt(d.comm$common.dispersion) [1] 0.2218773 > save(d.comm,file="CD14_dispersion_replicates_noTMM.RData") > sqrt(d.comm$common.dispersion) [1] 0.2208855 > save(d.comm,file="CD14_dispersion_replicates_noTMM_no_chrM.RData")
rDNA and chrM survey
Since Charles and a few people wondered if chrM mapped reads are correlated with rRNA reads, Kawaji plotted them to show their correlation. This support the strategy to remove chrM mapped reads from the library sizes.
Results
Methods
shell commands in ruby/Rake:
FileList["#{ROOTDIR}/UPDATE_012/f5pipeline/human.primary_cell.hCAGE/*.rdna.fa.gz"].each do |infile|
sh %! echo -ne $(basename #{infile} .nobarcode.rdna.fa.gz)"\t" >> #{t.name}!
sh %! tmp=$(grep -c '>' #{infile}); echo -ne "$tmp\t" >> #{t.name}!
infile = infile.sub(".nobarcode.rdna.fa.gz", ".ctss.bed.gz")
sh %! gunzip -c #{infile} | grep ^chrM | awk 'BEGIN{sum=0}{sum += $5}END{print sum}' >> #{t.name}!
end
