=item results *** results consistency: -- exon matches at basic level compare (blat > tandynear) for near exon dupls. show cele == dappulx, both 2x> dros -- but tandy find processing, nearly same, gives dappulx 1.5+> cele > dros ** what is difference in two? which is most correct? need to turn exon dupls into gene dupls to have biol. sense; -- can we use near-found exons & their gene models to make sense or are near-notfound required? ** add in full proteome blast compare (Dappux: NCBI_Gno; Cele: ref; Dros: ref) with near/far counts ** Repeats: repeatmasked daphnia genome has very few high-copy genes/exons; celegans has almost none (rep masked? or cleaned otherwise) drosophila are another story. Dmel is cleanest; but could use transposon repeat masking; Dere, Dmoj have per big scaffold some 400+ genes (over all predictors) with > 50 repeats (some with 400..1000 copies) Mostly these affect the Far category ** should adjust Near,Far to at least look at 1 intermediate (15kb, 30-50kb, >50kb) ** maybe 10kb, 20kb, 30kb, 40kb, .. # dpulex, all of clean genome perl $td/tandynear.perl -debug \ $em/daphe/scaffold_*/dpulex1_exons.nr $em/daphe/scaffold_*/dpulex1_exons.nr.blatf8 \ > & dpulex1-tandynear.txt & # exon_fasta_ids ref=scaffold_999 ngroup=3; ngenes=7; ntrans=7; nexons=17; naltexons=1 # exon_fasta_ids ref=scaffold_998 ngroup=1; ngenes=1; ntrans=1; nexons=2; naltexons=0 ... # exon_fasta_ids ref=scaffold_10 ngroup=3; ngenes=1366; ntrans=1430; nexons=6564; naltexons=4191 # exon_fasta_ids ref=scaffold_1 ngroup=3; ngenes=1707; ntrans=1892; nexons=10220; naltexons=6096 dpulex, Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside DP_DGIL_SNO_ 364605 264860 16676 10331 13969 52942 5827 DP_DGIL_SNO_ freq 1.000 0.063 0.039 0.053 0.200 0.022 DP_DGIL_SNO_ found 1.000 0.383 0.432 0.438 0.317 0.313 DP_DGIL_SNO_ anyfnd 1.000 0.431 0.473 0.481 0.358 0.333 DP_DGIL_SNO_ fnd/any 1.000 0.889 0.913 0.910 0.885 0.940 Dappu 226348 179398 8550 5633 7417 22284 3066 Dappu freq 1.000 0.048 0.031 0.041 0.124 0.017 Dappu found 1.000 0.271 0.309 0.348 0.328 0.159 Dappu anyfnd 1.000 0.441 0.470 0.508 0.465 0.350 Dappu fnd/any 1.000 0.614 0.658 0.685 0.706 0.456 NCBI_GNO_ 308944 231198 14151 8573 11664 39329 4029 NCBI_GNO_ freq 1.000 0.061 0.037 0.050 0.170 0.017 NCBI_GNO_ found 1.000 0.422 0.450 0.472 0.356 0.292 NCBI_GNO_ anyfnd 1.000 0.461 0.492 0.513 0.416 0.360 NCBI_GNO_ fnd/any 1.000 0.915 0.914 0.921 0.855 0.811 # dpulex, all of clean genome, >= 90% alignment Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside ## fixme: crossref missing # DP_DGIL_SNO_ 196384 161008 5699 3736 5356 18847 1738 # DP_DGIL_SNO_ freq 1.000 0.035 0.023 0.033 0.117 0.011 # DP_DGIL_SNO_ found 1.000 0.526 0.565 0.499 0.393 0.482 # DP_DGIL_SNO_ anyfnd 1.000 0.582 0.609 0.543 0.435 0.509 # DP_DGIL_SNO_ fnd/any 1.000 0.904 0.928 0.920 0.903 0.947 # # Dappu 153514 134424 3303 2221 3202 9299 1065 # Dappu freq 1.000 0.025 0.017 0.024 0.069 0.008 # Dappu found 1.000 0.450 0.479 0.488 0.473 0.285 # Dappu anyfnd 1.000 0.735 0.730 0.714 0.651 0.632 # Dappu fnd/any 1.000 0.612 0.656 0.684 0.726 0.450 # # NCBI_GNO_ 214521 177386 6393 4243 6054 18825 1620 # NCBI_GNO_ freq 1.000 0.036 0.024 0.034 0.106 0.009 # NCBI_GNO_ found 1.000 0.713 0.698 0.703 0.564 0.531 # NCBI_GNO_ anyfnd 1.000 0.770 0.758 0.754 0.634 0.649 # NCBI_GNO_ fnd/any 1.000 0.926 0.921 0.932 0.890 0.818 # # using MINALIGN=0.9 reduces near,far; increases found rates (to .7,.8) but # # ratio of group-found / any-found is ~same # perl $td/tandynear.perl -MINALIGN 0.9 $em/daphe/scaffold_?/dpulex1_exons.nr.blatf8 $em/daphe/scaffold_?/dpulex1_exons.nr #.................................................... # dmoj, Dros. mojavensis, tandynear stats # dmoj2, all euchromatin, subset predictors, perl $td/tandynear.perl -debug -skip dmoj1.all.dupids \ $em/dmoj2/scaffold_6*/dmoj_caf060210_exons.nr $em/dmoj2/scaffold_6*/dmoj_caf060210_exons.nr.blatf8 \ >& dmoj2-tandynear-nodups.txt & melon.% cat $em/dmoj2/dmoj2-tandynear-nodups.txt # exon_fasta_ids ref=scaffold_6680 ngroup=4; ngenes=12987; ntrans=13738; nexons=36091; naltexons=28390 # exon_fasta_ids ref=scaffold_6654 ngroup=4; ngenes=1571; ntrans=1688; nexons=4622; naltexons=4139 # exon_fasta_ids ref=scaffold_6541 ngroup=4; ngenes=1188; ntrans=1213; nexons=2154; naltexons=971 # exon_fasta_ids ref=scaffold_6540 ngroup=4; ngenes=18096; ntrans=19092; nexons=53273; naltexons=43650 # exon_fasta_ids ref=scaffold_6500 ngroup=4; ngenes=14912; ntrans=15450; nexons=39866; naltexons=31659 # exon_fasta_ids ref=scaffold_6498 ngroup=4; ngenes=1342; ntrans=1397; nexons=3647; naltexons=2264 # exon_fasta_ids ref=scaffold_6496 ngroup=4; ngenes=14180; ntrans=14903; nexons=43049; naltexons=36573 # exon_fasta_ids ref=scaffold_6482 ngroup=4; ngenes=1325; ntrans=1358; nexons=2941; naltexons=1876 # exon_fasta_ids ref=scaffold_6473 ngroup=4; ngenes=7766; ntrans=8233; nexons=20059; naltexons=13842 # exon_fasta_ids ref=scaffold_6359 ngroup=4; ngenes=2068; ntrans=2214; nexons=5830; naltexons=4080 # exon_fasta_ids ref=scaffold_6328 ngroup=4; ngenes=2153; ntrans=2318; nexons=6239; naltexons=5057 # exon_fasta_ids ref=scaffold_6308 ngroup=4; ngenes=1660; ntrans=1781; nexons=4385; naltexons=3231 # skipped ids n=2320 dmoj, Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside GI_BREN_NSC_ 191818 148234 3426 1547 2046 35195 1370 GI_BREN_NSC_ freq 1.000 0.023 0.010 0.014 0.237 0.009 GI_BREN_NSC_ found 1.000 0.320 0.363 0.203 0.034 0.247 GI_BREN_NSC_ anyfnd 1.000 0.383 0.416 0.241 0.053 0.296 GI_BREN_NSC_ fnd/any 1.000 0.835 0.874 0.840 0.639 0.837 GI_DGIL_SNO_ 284068 179198 4991 2234 3519 92934 1192 GI_DGIL_SNO_ freq 1.000 0.028 0.012 0.020 0.519 0.007 GI_DGIL_SNO_ found 1.000 0.311 0.354 0.191 0.062 0.260 GI_DGIL_SNO_ anyfnd 1.000 0.359 0.386 0.208 0.069 0.285 GI_DGIL_SNO_ fnd/any 1.000 0.866 0.916 0.917 0.890 0.912 GI_NCBI_GNO_ 163028 144004 3944 1401 1812 11198 669 * low far rate *might* be due to GI_NCBI_GNO_ freq 1.000 0.027 0.010 0.013 0.078 0.005 screening out known Transposons GI_NCBI_GNO_ found 1.000 0.388 0.453 0.241 0.072 0.217 compare others also to partial below GI_NCBI_GNO_ anyfnd 1.000 0.418 0.480 0.265 0.082 0.269 * likely scaffold-specific repeats GI_NCBI_GNO_ fnd/any 1.000 0.928 0.945 0.909 0.885 0.806 ? maybe add tandy's repeat skimmer? GLEAN_ 181168 142240 4029 1438 1818 30839 804 GLEAN_ freq 1.000 0.028 0.010 0.013 0.217 0.006 GLEAN_ found 1.000 0.366 0.427 0.209 0.042 0.216 GLEAN_ anyfnd 1.000 0.413 0.470 0.249 0.077 0.320 GLEAN_ fnd/any 1.000 0.886 0.908 0.839 0.548 0.677 # dmoj1, Dros. moj., subset euchromatin, all predictor perl $td/tandynear.perl -debug -skip dmoj1.all.dupids \ $em/dmoj1/scaffold_6*/dmoj_caf060210_exons.nr $em/dmoj1/scaffold_6*/dmoj_caf060210_exons.nr.blatf8 \ >& dmoj1-tandynear-nodups.txt & melon.% cat $em/dmoj1/dmoj1-tandynear-nodups.txt # exon_fasta_ids ref=scaffold_6680 ngroup=11; ngenes=27884; ntrans=38777; nexons=115848; naltexons=100979 # exon_fasta_ids ref=scaffold_6541 ngroup=11; ngenes=1875; ntrans=1996; nexons=3592; naltexons=2308 # exon_fasta_ids ref=scaffold_6473 ngroup=11; ngenes=15372; ntrans=20486; nexons=58949; naltexons=48780 # skipped ids n=1142232 dmoj1, Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside GI_BATZ_CNA_ 101656 88502 1022 284 628 10790 430 GI_BATZ_CNA_ freq 1.000 0.012 0.003 0.007 0.122 0.005 GI_BATZ_CNA_ found 1.000 0.266 0.415 0.115 0.022 0.435 GI_BATZ_CNA_ anyfnd 1.000 0.411 0.440 0.135 0.063 0.495 GI_BATZ_CNA_ fnd/any 1.000 0.648 0.944 0.847 0.349 0.878 GI_BREN_NSC_ 120066 104333 1366 375 746 12828 418 GI_BREN_NSC_ freq 1.000 0.013 0.004 0.007 0.123 0.004 GI_BREN_NSC_ found 1.000 0.316 0.357 0.101 0.023 0.428 GI_BREN_NSC_ anyfnd 1.000 0.428 0.381 0.141 0.053 0.550 GI_BREN_NSC_ fnd/any 1.000 0.740 0.937 0.714 0.427 0.778 GI_DGIL_SNO_ 143630 106561 1818 562 1202 32970 517 GI_DGIL_SNO_ freq 1.000 0.017 0.005 0.011 0.309 0.005 GI_DGIL_SNO_ found 1.000 0.338 0.324 0.116 0.040 0.321 GI_DGIL_SNO_ anyfnd 1.000 0.401 0.377 0.146 0.054 0.404 GI_DGIL_SNO_ fnd/any 1.000 0.842 0.858 0.800 0.742 0.794 GI_EISE_CEX_ 278661 266959 2182 325 951 6696 1548 GI_EISE_CEX_ freq 1.000 0.008 0.001 0.004 0.025 0.006 GI_EISE_CEX_ found 1.000 0.308 0.222 0.091 0.026 0.360 GI_EISE_CEX_ anyfnd 1.000 0.554 0.222 0.091 0.032 0.452 GI_EISE_CEX_ fnd/any 1.000 0.557 1.000 1.000 0.834 0.797 GI_EISE_CGW_ 97604 95121 785 122 257 716 603 GI_EISE_CGW_ freq 1.000 0.008 0.001 0.003 0.008 0.006 GI_EISE_CGW_ found 1.000 0.257 0.418 0.008 0.053 0.154 GI_EISE_CGW_ anyfnd 1.000 0.469 0.418 0.148 0.053 0.391 GI_EISE_CGW_ fnd/any 1.000 0.549 1.000 0.053 1.000 0.394 GI_NCBI_GNO_ 112838 104880 1826 315 614 4952 251 GI_NCBI_GNO_ freq 1.000 0.017 0.003 0.006 0.047 0.002 GI_NCBI_GNO_ found 1.000 0.464 0.302 0.119 0.022 0.386 GI_NCBI_GNO_ anyfnd 1.000 0.505 0.346 0.166 0.036 0.466 GI_NCBI_GNO_ fnd/any 1.000 0.918 0.872 0.716 0.594 0.829 GI_PACH_GMP_ 211179 208307 904 129 227 977 635 GI_PACH_GMP_ freq 1.000 0.004 0.001 0.001 0.005 0.003 GI_PACH_GMP_ found 1.000 0.021 0.000 0.000 0.094 0.000 GI_PACH_GMP_ anyfnd 1.000 0.488 0.310 0.000 0.094 0.485 GI_PACH_GMP_ fnd/any 1.000 0.043 0.000 0.000 1.000 0.000 GI_RGUI_GID_ 106999 90958 1033 294 522 13398 794 GI_RGUI_GID_ freq 1.000 0.011 0.003 0.006 0.147 0.009 GI_RGUI_GID_ found 1.000 0.274 0.340 0.023 0.029 0.378 GI_RGUI_GID_ anyfnd 1.000 0.379 0.381 0.071 0.057 0.453 GI_RGUI_GID_ fnd/any 1.000 0.724 0.893 0.324 0.506 0.833 GLEAN_ 127505 110774 1976 393 755 13390 217 GLEAN_ freq 1.000 0.018 0.004 0.007 0.121 0.002 GLEAN_ found 1.000 0.440 0.384 0.114 0.027 0.037 GLEAN_ anyfnd 1.000 0.505 0.448 0.171 0.067 0.244 GLEAN_ fnd/any 1.000 0.872 0.858 0.667 0.400 0.151 TRdmoj_ 102335 97705 1214 202 366 2300 548 TRdmoj_ freq 1.000 0.012 0.002 0.004 0.024 0.006 TRdmoj_ found 1.000 0.523 0.322 0.186 0.144 0.155 TRdmoj_ anyfnd 1.000 0.555 0.322 0.216 0.145 0.332 TRdmoj_ fnd/any 1.000 0.942 1.000 0.861 0.991 0.467 dmoj_GLEANR_ 121089 105286 1827 383 721 12658 214 dmoj_GLEANR_ freq 1.000 0.017 0.004 0.007 0.120 0.002 dmoj_GLEANR_ found 1.000 0.453 0.334 0.107 0.024 0.000 dmoj_GLEANR_ anyfnd 1.000 0.503 0.392 0.173 0.064 0.266 dmoj_GLEANR_ fnd/any 1.000 0.900 0.853 0.616 0.374 0.000 # # GI_EISE_CGW_ 43623 42371 441 540 271 ** removed duplicate exons (by id) # GI_EISE_CGW_ freq 1.000 0.010 0.013 0.006 now consistent w/ similar predictors (exonerate, dpulex gw) # GI_EISE_CGW_ found 1.000 0.279 0.033 0.151 # GI_EISE_CGW_ anyfnd 1.000 0.456 0.061 0.395 # GI_EISE_CGW_ fnd/any 1.000 0.612 0.545 0.383 # # GI_NCBI_GNO_ 68740 63979 1259 3341 161 # GI_NCBI_GNO_ freq 1.000 0.020 0.052 0.003 # GI_NCBI_GNO_ found 1.000 0.442 0.032 0.435 # GI_NCBI_GNO_ anyfnd 1.000 0.496 0.060 0.491 # GI_NCBI_GNO_ fnd/any 1.000 0.891 0.543 0.886 # **^ cmp to EISE_CGW; scaffold_6680 has NCBI_GNO 9489 exons; 8645 primary in exons.nr; # ** 42125 NCBI exons.nr alt ids, excluding partial matches (about same as EISE_CGW) # # # GI_EISE_CGW_ 154467 150220 1344 1887 1016 ** highest exon count; # # GI_EISE_CGW_ freq 1.000 0.009 0.013 0.007 are these really alt genes, not alt exons/gene? # # GI_EISE_CGW_ found 1.000 0.365 0.144 0.322 # # GI_EISE_CGW_ anyfnd 1.000 0.476 0.144 0.531 # # GI_EISE_CGW_ fnd/any 1.000 0.766 0.996 0.606 #.............. ** dmoj2:GI_EISE_CGW_ ^^ odd problem: scaffold_6680 has only 19171 EISE_CGW exons; but ** $em/dmoj1/scaffold_6680/dmoj_caf060210_exons.nr has 55231 entries: partial matches? ** $em/dmoj1/scaffold_6680/ _exons.fa has 19171 EISE_CGW exons : ok ** removing pm, partials, get 41322 exons (alt matches to other predictor, same loc) ** could this be causing some kind of double/extra counting? ** maybe exon overlaps? ** YES, for diff gene/tr ids: alt transcripts; ** some are ident locs; some are overlaps ** removing dupl. exons good; do others with high alt-exon counts need this? all to be fair? gzgrep '^>GI_EISE_CGW' $em/dmoj1/scaffold_6680/dmoj_caf060210_exons.fa.gz | perl -ne'($d)=m/^>(\S+)/; ($ d,$x,$b,$e)=split(/[\.:-]/,$d); print join("\t",$b,$e,$d,$x),"\n";' | sort -k1,1n -k2,2nr -k3,3 -k4,4n | more ** should weed out these alt exons/alt tr from GFF before running blat-tandy? gzcat $em/dmoj1/scaffold_6*/dmoj_caf060210_exons.fa.gz | grep '^>GI_EISE_CGW' | wc 28432 gzcat $em/dmoj1/scaffold_6*/dmoj_caf060210_exons.fa.gz | grep '^>GI_EISE_CGW' | perl -ne'($d)=m/^>(\S+)/ ; ($d,$x,$b,$e)=split(/[\.:-]/,$d); ($r)=m/loc=(\w+)/; print join("\t",$b,$e,$r,$d,$x),"\n";' | sort -k3,3 -k1,1 n -k2,2nr -k4,4 -k5,5n | perl -ne'($b,$e,$r,$d,$x)=split; print "$d.$x\ni"if($r eq $lr && $b <= $le && $e >= $lb) -k2,2nr -k4,4 -k5,5n | perl -ne'($b,$e,$r,$d,$x)=split; print "$d.$x\n" if($r eq $lr && $b <= $le && $e >= $lb ); ($lr,$lb,$le)= ($r,$b,$e);' > dmoj1.EISE_CGW.dupids wc dmoj1.EISE_CGW.dupids 16681 perl $td/tandynear.perl -skip dmoj1.EISE_CGW.dupids \ $em/dmoj1/scaffold_6680/dmoj_caf060210_exons.nr $em/dmoj1/scaffold_6680/dmoj_caf060210_exons.nr.blatf8 compare like treatment for GI_NCBI_GNO_, dmoj_GLEANR_ : have lower exon counts: GI_NCBI_GNO_: 14843, dmoj_GLEANR_: 14041 versus GI_EISE_CGW: 28432 gzcat $em/dmoj1/scaffold_6*/dmoj_caf060210_exons.fa.gz | grep '^>dmoj_GLEANR_' | perl -ne\ '($d)=m/^>(\S+)/; ($d,$x,$b,$e)=split(/[\.:-]/,$d); ($r)=m/loc=(\w+)/; print join("\t",$b,$e,$r,$d,$x),"\n";' \ | sort -k3,3 -k1,1n -k2,2nr -k4,4 -k5,5n \ | perl -ne'($b,$e,$r,$d,$x)=split; print "$d.$x\n" if($r eq $lr && $b <= $le && $e >= $lb); ($lr,$lb,$le)= ($r,$b,$e);' \ > dmoj1.dmoj_GLEANR.dupids 16681 16681 384397 dmoj1.EISE_CGW.dupids 218 218 5022 dmoj1.GI_NCBI_GNO.dupids 3278 3278 49105 dmoj1.TRdmoj.dupids << maybe a problem 12 12 240 dmoj1.dmoj_GLEANR.dupids #.............. #.................................................... # dere, Dros. erecta # full euchromatin genome, all predicts (removing alt-tr/exons: CEX,CGW) perl $td/tandynear.perl -debug \ $em/dere1/scaffold_*/dere_caf060210_exons.nr $em/dere1/scaffold_*/dere_caf060210_exons.nr.blatf8 \ > & $em/dere1/dere1-tandynear.txt & melon.% cat $em/dere1/dere1-tandynear.txt # exon_fasta_ids ref=scaffold_4929 ngroup=11; ngenes=35353; ntrans=37481; nexons=109762; naltexons=98761 # exon_fasta_ids ref=scaffold_4845 ngroup=11; ngenes=33324; ntrans=35574; nexons=104795; naltexons=95193 # exon_fasta_ids ref=scaffold_4820 ngroup=11; ngenes=15685; ntrans=16784; nexons=52873; naltexons=48472 # exon_fasta_ids ref=scaffold_4784 ngroup=11; ngenes=34803; ntrans=36958; nexons=104825; naltexons=93668 # exon_fasta_ids ref=scaffold_4770 ngroup=11; ngenes=26551; ntrans=28227; nexons=84087; naltexons=77008 # exon_fasta_ids ref=scaffold_4690 ngroup=11; ngenes=24285; ntrans=25918; nexons=74880; naltexons=67855 # exon_fasta_ids ref=scaffold_4644 ngroup=11; ngenes=3633; ntrans=3863; nexons=10865; naltexons=9811 # exon_fasta_ids ref=scaffold_4512 ngroup=11; ngenes=889; ntrans=974; nexons=4115; naltexons=3957 dere, Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside GG_BATZ_CNA_ 398441 349660 6332 3774 5699 30955 2021 GG_BATZ_CNA_ freq 1.000 0.018 0.011 0.016 0.089 0.006 GG_BATZ_CNA_ found 1.000 0.505 0.694 0.683 0.052 0.242 GG_BATZ_CNA_ anyfnd 1.000 0.574 0.756 0.785 0.096 0.411 GG_BATZ_CNA_ fnd/any 1.000 0.878 0.918 0.871 0.547 0.589 GG_BREN_NSC_ 458292 391703 6580 4095 6228 47058 2628 GG_BREN_NSC_ freq 1.000 0.017 0.010 0.016 0.120 0.007 GG_BREN_NSC_ found 1.000 0.502 0.687 0.680 0.041 0.292 GG_BREN_NSC_ anyfnd 1.000 0.561 0.753 0.749 0.093 0.441 GG_BREN_NSC_ fnd/any 1.000 0.894 0.912 0.907 0.446 0.662 GG_DGIL_SNO_ 475651 366638 8911 5661 8935 83189 2317 GG_DGIL_SNO_ freq 1.000 0.024 0.015 0.024 0.227 0.006 GG_DGIL_SNO_ found 1.000 0.517 0.713 0.656 0.090 0.353 GG_DGIL_SNO_ anyfnd 1.000 0.567 0.759 0.738 0.111 0.438 GG_DGIL_SNO_ fnd/any 1.000 0.912 0.938 0.888 0.813 0.807 GG_EISE_CEX_ 360416 350618 4193 806 668 3093 1038 GG_EISE_CEX_ freq 1.000 0.012 0.002 0.002 0.009 0.003 GG_EISE_CEX_ found 1.000 0.181 0.156 0.208 0.057 0.137 GG_EISE_CEX_ anyfnd 1.000 0.364 0.270 0.238 0.074 0.352 GG_EISE_CEX_ fnd/any 1.000 0.498 0.578 0.874 0.770 0.389 GG_EISE_CGW_ 376454 367285 3934 868 668 2352 1347 GG_EISE_CGW_ freq 1.000 0.011 0.002 0.002 0.006 0.004 GG_EISE_CGW_ found 1.000 0.154 0.245 0.195 0.091 0.123 GG_EISE_CGW_ anyfnd 1.000 0.356 0.346 0.262 0.112 0.306 GG_EISE_CGW_ fnd/any 1.000 0.432 0.710 0.743 0.814 0.403 GG_NCBI_GNO_ 431080 395008 7870 3320 4019 19449 1414 GG_NCBI_GNO_ freq 1.000 0.020 0.008 0.010 0.049 0.004 GG_NCBI_GNO_ found 1.000 0.472 0.565 0.592 0.032 0.115 GG_NCBI_GNO_ anyfnd 1.000 0.514 0.648 0.675 0.053 0.203 GG_NCBI_GNO_ fnd/any 1.000 0.919 0.873 0.877 0.616 0.568 GG_PACH_GMP_ 407448 391327 5644 3135 4443 2204 695 GG_PACH_GMP_ freq 1.000 0.014 0.008 0.011 0.006 0.002 GG_PACH_GMP_ found 1.000 0.236 0.453 0.422 0.084 0.055 GG_PACH_GMP_ anyfnd 1.000 0.508 0.755 0.839 0.164 0.271 GG_PACH_GMP_ fnd/any 1.000 0.464 0.600 0.503 0.511 0.202 GG_RGUI_GID_ 401161 345024 6156 4729 7292 34282 3678 GG_RGUI_GID_ freq 1.000 0.018 0.014 0.021 0.099 0.011 GG_RGUI_GID_ found 1.000 0.537 0.774 0.777 0.048 0.344 GG_RGUI_GID_ anyfnd 1.000 0.593 0.818 0.830 0.081 0.455 GG_RGUI_GID_ fnd/any 1.000 0.906 0.946 0.936 0.594 0.755 GLEAN_ 482149 411244 9121 5564 8206 46200 1814 GLEAN_ freq 1.000 0.022 0.014 0.020 0.112 0.004 GLEAN_ found 1.000 0.530 0.711 0.704 0.050 0.346 GLEAN_ anyfnd 1.000 0.588 0.773 0.803 0.116 0.514 GLEAN_ fnd/any 1.000 0.902 0.920 0.876 0.434 0.673 TRdere_ 438817 392960 7975 5220 7647 23621 1394 TRdere_ freq 1.000 0.020 0.013 0.019 0.060 0.004 TRdere_ found 1.000 0.567 0.802 0.857 0.321 0.122 TRdere_ anyfnd 1.000 0.605 0.818 0.861 0.324 0.321 TRdere_ fnd/any 1.000 0.937 0.981 0.996 0.991 0.379 dere_GLEANR_ 448419 385918 8457 5361 7985 37962 2736 dere_GLEANR_ freq 1.000 0.022 0.014 0.021 0.098 0.007 dere_GLEANR_ found 1.000 0.543 0.724 0.717 0.040 0.163 dere_GLEANR_ anyfnd 1.000 0.608 0.791 0.816 0.106 0.298 dere_GLEANR_ fnd/any 1.000 0.892 0.916 0.878 0.374 0.545 ## vvvv Fixed; needed altid crossref # ** Same count is oddly increasing lexically: is this artifact of tandy reduce # step combining same exons by class?? or other artifact? Input exon counts/group, Dere1 melon.% grep '^>' scaffold_4690/dere_caf060210_exons.fa | perl -ne'm/>(\D+)/;print "$1\n";' | sort | uniq -c 6999 GG_BATZ_CNA_ 7742 GG_BREN_NSC_ 10993 GG_DGIL_SNO_ 7163 GG_EISE_CEX_ 7188 GG_EISE_CGW_ 7813 GG_NCBI_GNO_ 7065 GG_PACH_GMP_ 7580 GG_RGUI_GID_mRNA_ 7809 GLEAN_ 7177 TRdere_ 7596 dere_GLEANR_ Reduced exons, primary ID only -- no big diff, no lexical drift? melon.% grep '^>' scaffold_4690/dere_caf060210_exons.nr | perl -ne'm/>(\D+)/;print "$1\n";' | sort | uniq -c 6816 GG_BATZ_CNA_ 7476 GG_BREN_NSC_ 10747 GG_DGIL_SNO_ 7027 GG_EISE_CEX_ 6903 GG_EISE_CGW_ 7180 GG_NCBI_GNO_ 6923 GG_PACH_GMP_ 7245 GG_RGUI_GID_mRNA_ 6638 GLEAN_ 4082 TRdere_ 3843 dere_GLEANR_ # with altids: ** HERE IS PROBLEM; some lexical increase with ID class in altids ** grep '^>' scaffold_4690/dere_caf060210_exons.nr | perl -ne'm/>(\D+)/; $d=$1; ($a)=m/altids=(\S+)/; $a= ~s/pm,.*$//; @a=split",",$a; map{s/\d.*$//;}@a; print join("\n",$d,@a),"\n";' | sort | uniq -c 12856 GG_BATZ_CNA_ 19195 GG_BREN_NSC_ 24273 GG_DGIL_SNO_ 23656 GG_EISE_CEX_ 29155 GG_EISE_CGW_ 36533 GG_NCBI_GNO_ 37454 GG_PACH_GMP_ 40102 GG_RGUI_GID_mRNA_ 51171 GLEAN_ 52762 TRdere_ 52501 dere_GLEANR_ 12856 GG_BATZ_CNA_ 19195 GG_BREN_NSC_ .. 52762 TRdere_ 52501 dere_GLEANR_ # # ** partial data # melon.% more $em/dere1/dere1-tandynear-part.txt # # exon_fasta_ids ref=scaffold_4784 ngroup=11; ngenes=34755; ntrans=36906; nexons=104825; naltexons=334370 # # exon_fasta_ids ref=scaffold_4690 ngroup=11; ngenes=24253; ntrans=25881; nexons=74880; naltexons=245642 # # exon_fasta_ids ref=scaffold_4644 ngroup=11; ngenes=3628; ntrans=3858; nexons=10865; naltexons=35030 # # exon_fasta_ids ref=scaffold_4512 ngroup=11; ngenes=886; ntrans=971; nexons=4115; naltexons=10797 # # Tandy exon match types per predictor group # Group Total Same Near8k Near16k Near48k Far Inside # GG_BATZ_CNA_ 20707 16929 224 57 122 3301 74 # GG_BATZ_CNA_ freq 1.000 0.013 0.003 0.007 0.195 0.004 # GG_BATZ_CNA_ found 1.000 0.299 0.088 0.090 0.092 0.068 # GG_BATZ_CNA_ anyfnd 1.000 0.362 0.123 0.180 0.127 0.149 # GG_BATZ_CNA_ fnd/any 1.000 0.827 0.714 0.500 0.728 0.455 # # GG_BREN_NSC_ 38924 32441 345 114 229 5645 150 # GG_BREN_NSC_ freq 1.000 0.011 0.004 0.007 0.174 0.005 # GG_BREN_NSC_ found 1.000 0.380 0.175 0.083 0.041 0.147 # GG_BREN_NSC_ anyfnd 1.000 0.458 0.298 0.175 0.090 0.180 # GG_BREN_NSC_ fnd/any 1.000 0.829 0.588 0.475 0.461 0.815 # # GG_DGIL_SNO_ 58288 43191 525 136 421 13513 502 # GG_DGIL_SNO_ freq 1.000 0.012 0.003 0.010 0.313 0.012 # GG_DGIL_SNO_ found 1.000 0.309 0.169 0.152 0.108 0.406 # GG_DGIL_SNO_ anyfnd 1.000 0.373 0.338 0.183 0.128 0.414 # GG_DGIL_SNO_ fnd/any 1.000 0.827 0.500 0.831 0.844 0.981 # # GG_EISE_CEX_ 47988 46444 468 79 129 706 162 # GG_EISE_CEX_ freq 1.000 0.010 0.002 0.003 0.015 0.003 # GG_EISE_CEX_ found 1.000 0.152 0.101 0.140 0.020 0.130 # GG_EISE_CEX_ anyfnd 1.000 0.335 0.101 0.147 0.027 0.191 # GG_EISE_CEX_ fnd/any 1.000 0.452 1.000 0.947 0.737 0.677 # # GG_EISE_CGW_ 62019 60827 463 91 122 398 118 # GG_EISE_CGW_ freq 1.000 0.008 0.001 0.002 0.007 0.002 # GG_EISE_CGW_ found 1.000 0.140 0.143 0.164 0.068 0.153 # GG_EISE_CGW_ anyfnd 1.000 0.361 0.143 0.213 0.085 0.246 # GG_EISE_CGW_ fnd/any 1.000 0.389 1.000 0.769 0.794 0.621 # # GG_NCBI_GNO_ 82413 78859 923 236 247 1726 422 # GG_NCBI_GNO_ freq 1.000 0.012 0.003 0.003 0.022 0.005 # GG_NCBI_GNO_ found 1.000 0.335 0.165 0.134 0.045 0.109 # GG_NCBI_GNO_ anyfnd 1.000 0.389 0.178 0.182 0.076 0.154 # GG_NCBI_GNO_ fnd/any 1.000 0.861 0.929 0.733 0.583 0.708 # # GG_PACH_GMP_ 81565 80399 431 85 167 363 120 # GG_PACH_GMP_ freq 1.000 0.005 0.001 0.002 0.005 0.001 # GG_PACH_GMP_ found 1.000 0.074 0.000 0.204 0.052 0.000 # GG_PACH_GMP_ anyfnd 1.000 0.399 0.000 0.240 0.105 0.308 # GG_PACH_GMP_ fnd/any 1.000 0.186 0.000 0.850 0.500 0.000 # # GG_RGUI_GID_ 90857 83796 562 130 315 5590 464 # GG_RGUI_GID_ freq 1.000 0.007 0.002 0.004 0.067 0.006 # GG_RGUI_GID_ found 1.000 0.342 0.131 0.137 0.067 0.377 # GG_RGUI_GID_ anyfnd 1.000 0.409 0.223 0.187 0.105 0.485 # GG_RGUI_GID_ fnd/any 1.000 0.835 0.586 0.729 0.644 0.778 # # GLEAN_ 128593 115241 1134 264 484 11174 296 # GLEAN_ freq 1.000 0.010 0.002 0.004 0.097 0.003 # GLEAN_ found 1.000 0.422 0.216 0.163 0.064 0.220 # GLEAN_ anyfnd 1.000 0.487 0.254 0.256 0.156 0.314 # GLEAN_ fnd/any 1.000 0.866 0.851 0.637 0.411 0.699 # # TRdere_ 135436 127854 926 177 373 5837 269 # TRdere_ freq 1.000 0.007 0.001 0.003 0.046 0.002 # TRdere_ found 1.000 0.382 0.045 0.214 0.343 0.175 # TRdere_ anyfnd 1.000 0.445 0.045 0.236 0.346 0.409 # TRdere_ fnd/any 1.000 0.859 1.000 0.909 0.993 0.427 # # dere_GLEANR_ 137409 126447 1084 256 507 8849 266 # dere_GLEANR_ freq 1.000 0.009 0.002 0.004 0.070 0.002 # dere_GLEANR_ found 1.000 0.423 0.258 0.140 0.046 0.143 # dere_GLEANR_ anyfnd 1.000 0.519 0.305 0.237 0.139 0.398 # dere_GLEANR_ fnd/any 1.000 0.815 0.846 0.592 0.331 0.358 #.................................................... # dmel, Dros. mel. -- redo tandy w/ all CAF1 predictors # ** includes alt splice exons?; :( ids are all FBtr transcript ; cant tell gene span from this # should reframe _isnear to use gene-span, distinguish inside/outside of genespan perl $td/tandynear.perl -debug -skip dmel4.ncrna-exon.dupids -nogroup \ $em/dmel4/{2,3,4,X,U}*/dmel4_exons.nr.blatf8 $em/dmel4/{2,3,4,X,U}*/dmel4_exons.nr \ > & dmel4-tandynear-nodups.txt & # exon_fasta_ids ref=X ngroup=1; ngenes=2058; ntrans=3138; nexons=10009; naltexons=170 # exon_fasta_ids ref=4 ngroup=1; ngenes=89; ntrans=205; nexons=875; naltexons=16 # exon_fasta_ids ref=3R ngroup=1; ngenes=3060; ntrans=4753; nexons=15996; naltexons=151 # exon_fasta_ids ref=3L ngroup=1; ngenes=2376; ntrans=3737; nexons=12089; naltexons=138 # exon_fasta_ids ref=2R ngroup=1; ngenes=2524; ntrans=4057; nexons=13336; naltexons=264 # exon_fasta_ids ref=2L ngroup=1; ngenes=2332; ntrans=3565; nexons=11483; naltexons=216 # skipped ids n=81418 Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside all 85089 63295 3501 3445 7384 7078 386 all freq 1.000 0.055 0.054 0.117 0.112 0.006 all found 1.000 0.693 0.900 0.941 0.852 0.070 # Dros.mel, -minalign 0.9 Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside all 83547 63295 2909 3278 7270 6641 154 all freq 1.000 0.046 0.052 0.115 0.105 0.002 all found 1.000 0.832 0.944 0.956 0.907 0.175 # ncRNA high dupl. rate artifact ......... # Group Total Same Near Far Inside # all 113796 65000 24650 23553 593 # all freq 1.000 0.379 0.362 0.009 * odd high near artifact of ncRNA matches # all found 1.000 0.951 0.948 0.567 # * non-coding exons/genes are included in dmel4 annots, some high-dupl trna/ncrna # * 2R has 145% Near > Same ** ; 2L has 20% Near; X has 7% Near; 3L,3R have 4% Near/Same; ?????? # ^^^ this is the problem; 2R has some 2200 matches to 280 ncRNA genes; # 3L has 465 matches to 100 ncRNA genes; 2L has 911 matches to 138 ncRNA #............................................ # Dros. mel with predictions, removing ncRNA, removing duplicate exons (alt transcripts) # ** Maybe should be using transposon-masked genome fasta ** # are high predictor dupls (near, esp far) due to those?? ${dpid}_transposon.gff :: TE regions are high predict matches is there any way to guess if matches are TE regions? do multi-exon genes have introns in te-region matches? Try Dmel2R test case (TE rich) #...... perl $td/tandynear.perl -debug -skip dmel4c.ncrna-exon.dupids -group \ $em/dmel4c/{2,3,4,X}*/dmel4c_exons.nr.blatf8 $em/dmel4c/{2,3,4,X}*/dmel4c_exons.nr \ > & ! dmel4c-tandynear-nodups.txt & # exon_fasta_ids ref=X ngroup=6; ngenes=14947; ntrans=18297; nexons=56558; naltexons=130106 # exon_fasta_ids ref=4 ngroup=6; ngenes=498; ntrans=826; nexons=4433; naltexons=14567 # exon_fasta_ids ref=3R ngroup=6; ngenes=20936; ntrans=26302; nexons=87889; naltexons=210178 # exon_fasta_ids ref=3L ngroup=6; ngenes=16950; ntrans=21002; nexons=65754; naltexons=144495 # exon_fasta_ids ref=2R ngroup=6; ngenes=16938; ntrans=21605; nexons=73374; naltexons=202830 # exon_fasta_ids ref=2L ngroup=7; ngenes=15885; ntrans=19588; nexons=61867; naltexons=149745 # skipped ids n=449455 ** Far is high in predicteds due in part to Transposon matches Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside ## fixme: id crossref # CG_BATZ_CON_ 75248 52936 2782 2979 7493 8644 414 # CG_BATZ_CON_ freq 1.000 0.053 0.056 0.142 0.163 0.008 # CG_BATZ_CON_ found 1.000 0.768 0.933 0.974 0.931 0.362 # CG_BATZ_CON_ anyfnd 1.000 0.783 0.942 0.981 0.936 0.560 # CG_BATZ_CON_ fnd/any 1.000 0.982 0.990 0.993 0.994 0.647 # # CG_DGIL_SNO_ 136406 98413 3848 3504 8128 21585 928 # CG_DGIL_SNO_ freq 1.000 0.039 0.036 0.083 0.219 0.009 # CG_DGIL_SNO_ found 1.000 0.719 0.899 0.950 0.594 0.462 # CG_DGIL_SNO_ anyfnd 1.000 0.735 0.901 0.953 0.601 0.613 # CG_DGIL_SNO_ fnd/any 1.000 0.979 0.997 0.997 0.988 0.754 # # CG_NCBI_GNO_ 254883 181388 10431 10679 24411 27446 528 # CG_NCBI_GNO_ freq 1.000 0.058 0.059 0.135 0.151 0.003 # CG_NCBI_GNO_ found 1.000 0.805 0.906 0.971 0.867 0.320 # CG_NCBI_GNO_ anyfnd 1.000 0.838 0.915 0.981 0.897 0.335 # CG_NCBI_GNO_ fnd/any 1.000 0.960 0.990 0.991 0.966 0.955 # # CG_RGUI_GID_ 211624 147494 7048 8136 19768 26197 2981 # CG_RGUI_GID_ freq 1.000 0.048 0.055 0.134 0.178 0.020 # CG_RGUI_GID_ found 1.000 0.820 0.937 0.962 0.848 0.784 # CG_RGUI_GID_ anyfnd 1.000 0.850 0.952 0.976 0.876 0.813 # CG_RGUI_GID_ fnd/any 1.000 0.966 0.984 0.986 0.968 0.963 # # FBtr 193434 175278 5085 2890 4846 4868 467 # FBtr freq 1.000 0.029 0.016 0.028 0.028 0.003 # FBtr found 1.000 0.614 0.794 0.905 0.694 0.251 # FBtr anyfnd 1.000 0.682 0.813 0.912 0.700 0.272 # FBtr fnd/any 1.000 0.900 0.977 0.992 0.991 0.921 # # TRdmel_CG 377328 354917 8278 3823 4541 5170 599 # TRdmel_CG freq 1.000 0.023 0.011 0.013 0.015 0.002 # TRdmel_CG found 1.000 0.623 0.445 0.671 0.057 0.244 # TRdmel_CG anyfnd 1.000 0.718 0.586 0.866 0.575 0.277 # TRdmel_CG fnd/any 1.000 0.868 0.760 0.775 0.099 0.880 #............................................ # dmel4d, Dros. mel with predictions, masking Transposons, removing ncRNA, removing duplicate exons (alt transcripts) # ${dpid}_transposon.gff :: TE regions are high predict matches perl $td/tandynear.perl -debug -group \ $em/dmel4d/{2,3,4,X}*/dmel4c_exons.nr.blatf8 $em/dmel4d/{2,3,4,X}*/dmel4c_exons.nr \ > & dmel4d-tandynear.txt & # exon_fasta_ids ref=X ngroup=6; ngenes=15225; ntrans=16404; nexons=46692; naltexons=39881 # exon_fasta_ids ref=4 ngroup=6; ngenes=466; ntrans=534; nexons=2319; naltexons=2124 # exon_fasta_ids ref=3R ngroup=6; ngenes=21688; ntrans=23673; nexons=72875; naltexons=63588 # exon_fasta_ids ref=3L ngroup=6; ngenes=17274; ntrans=18780; nexons=54605; naltexons=46722 # exon_fasta_ids ref=2R ngroup=6; ngenes=17229; ntrans=18881; nexons=58764; naltexons=52311 # exon_fasta_ids ref=2L ngroup=7; ngenes=16282; ntrans=17649; nexons=51500; naltexons=44750 dmel, Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside CG_BATZ_CON_ 307047 223440 11670 11810 28039 30152 1936 CG_BATZ_CON_ freq 1.000 0.052 0.053 0.125 0.135 0.009 CG_BATZ_CON_ found 1.000 0.798 0.935 0.975 0.942 0.454 CG_BATZ_CON_ anyfnd 1.000 0.813 0.948 0.985 0.948 0.759 CG_BATZ_CON_ fnd/any 1.000 0.982 0.987 0.990 0.993 0.598 CG_DGIL_SNO_ 299329 216155 11338 11793 27593 30552 1898 CG_DGIL_SNO_ freq 1.000 0.052 0.055 0.128 0.141 0.009 CG_DGIL_SNO_ found 1.000 0.799 0.940 0.973 0.901 0.429 CG_DGIL_SNO_ anyfnd 1.000 0.810 0.941 0.976 0.912 0.676 CG_DGIL_SNO_ fnd/any 1.000 0.986 0.999 0.997 0.988 0.634 CG_NCBI_GNO_ 304359 221820 13302 12004 27215 29286 732 CG_NCBI_GNO_ freq 1.000 0.060 0.054 0.123 0.132 0.003 CG_NCBI_GNO_ found 1.000 0.767 0.857 0.913 0.868 0.186 CG_NCBI_GNO_ anyfnd 1.000 0.820 0.917 0.962 0.923 0.235 CG_NCBI_GNO_ fnd/any 1.000 0.935 0.934 0.950 0.941 0.791 CG_RGUI_GID_ 271281 190767 9629 11165 26856 29144 3720 CG_RGUI_GID_ freq 1.000 0.050 0.059 0.141 0.153 0.020 CG_RGUI_GID_ found 1.000 0.822 0.941 0.966 0.943 0.760 CG_RGUI_GID_ anyfnd 1.000 0.849 0.953 0.976 0.944 0.787 CG_RGUI_GID_ fnd/any 1.000 0.969 0.988 0.989 0.999 0.965 FBtr 212936 185476 6633 4371 8076 7897 483 FBtr freq 1.000 0.036 0.024 0.044 0.043 0.003 FBtr found 1.000 0.628 0.825 0.919 0.781 0.236 FBtr anyfnd 1.000 0.701 0.854 0.931 0.787 0.253 FBtr fnd/any 1.000 0.896 0.967 0.987 0.992 0.934 TRdmel_CG 222314 207143 6406 2139 1956 4181 489 TRdmel_CG freq 1.000 0.031 0.010 0.009 0.020 0.002 TRdmel_CG found 1.000 0.547 0.455 0.372 0.065 0.094 TRdmel_CG anyfnd 1.000 0.660 0.710 0.833 0.683 0.102 TRdmel_CG fnd/any 1.000 0.829 0.642 0.447 0.096 0.920 #.................................. # ///////// # FBtr 116239 69939 23872 21948 480 ** ncrna exons mistake # FBtr freq 1.000 0.341 0.314 0.007 ** odd; adding 2R jumps near from 8% to 34% # FBtr found 1.000 0.957 0.948 0.752 **?? have we left in ncrna ?? ARGH yes, 294 of them * how?? #........... ** Check these Dmel predictions for alt-tr/alt-exon dups: Group Total Same Near Far Inside CG_NCBI_GNO_ 102685 39590 10985 51843 267 (higher than DGIL_SNO looks like alt tr) CG_NCBI_GNO_ freq 1.000 0.277 1.309 0.007 * high rate of near maybe spurious? CG_RGUI_GID_mRNA_ 79542 27354 8770 43259 159 CG_RGUI_GID_mRNA_ freq 1.000 0.321 1.581 0.006 CG_DGIL_SNO_ 40370 18138 3531 18636 65 CG_DGIL_SNO_ freq 1.000 0.195 1.027 0.004 FBtr 43607 31780 2814 8911 102 FBtr freq 1.000 0.089 0.280 0.003 * consistent freq TRdmel_CG 84851 78482 1889 4241 239 TRdmel_CG freq 1.000 0.024 0.054 0.003 CG_BATZ_CON_ 29726 10071 3367 16244 44 CG_BATZ_CON_ freq 1.000 0.334 1.613 0.004 cat $em/dmel4c/*/dmel4c_exons.fa | grep '^>CG_BATZ_CON' | perl -ne\ '($d)=m/^>(\S+)/; ($d,$x,$b,$e)=split(/[\.:-]/,$d); ($r)=m/loc=(\w+)/; print join("\t",$b,$e,$r,$d,$x),"\n";' \ | sort -k3,3 -k1,1n -k2,2nr -k4,4 -k5,5n \ | perl -ne'($b,$e,$r,$d,$x)=split; print "$d.$x\n" if($r eq $lr && $b <= $le && $e >= $lb); ($lr,$lb,$le)= ($r,$b,$e);' \ > ! dmel4c.CG_BATZ_CON.dupids melon.% wc *dupids 0 0 0 dmel4c.CG_BATZ_CON.dupids 179 179 4120 dmel4c.CG_DGIL_SNO.dupids 9990 9990 230737 dmel4c.CG_NCBI_GNO.dupids << problem 0 0 0 dmel4c.CG_RGUI_GID.dupids 2305 2305 32368 dmel4c.FBtr.dupids << problem? 12909 12909 252699 dmel4c.TRdmel_CG.dupids << problem >> dmel4c.exon.dupids #............... # ** Maybe should be using transposon-masked genome fasta ** # are high predictor dupls (near, esp far) due to those?? ${dpid}_transposon.gff :: can be only 20 - 50 bp long; some are gene sized cat $em/dmel4c/2R/dmel4c_exons.nr.blatf8 | perl -ne \ 'BEGIN{ open(F,"dmel4c_transposon.gff"); while(){ chomp; ($r,$s,$t,$b,$e)=split; \ $bin=int(($b+$e)/2000); push @{$tpos{$r}{$bin}}, [$b,$e,$r,];} close(F);} \ chomp; my @v=split"\t"; \ my($qid1,$ref,$pctid,$alen,$tb,$te,$eval,$bits)=@v[0,1,2,3,8,9,10,11]; \ $bin=int(($tb+$te)/2000); my($qid,$qb,$qe)=split(/[:-]/,$qid1); \ $palign= $alen / 1+abs($qe-$qb); next unless($palign > 0.5); \ @tp= $tpos{$ref}{$bin} ? @{$tpos{$ref}{$bin}} : ();\ $olap=0; ($ob,$oe)=(0,0); foreach $tp (@tp) { \ if($tb < $$tp[1] && $te > $$tp[0]){ $olap++; ($ob,$oe)=($$tp[0],$$tp[1]); last;} }\ print "TEmatch: ",join(", ",$qid1,$ref,$tb,$te,$eval,$bits),"\n" if($olap); ' \ >> check predict ids for exon/intron structure in TE regions : same as main gene if main is not in TE, and if it is? # many TE match, 3+ exons, main predict is TE TEmatch: CG_NCBI_GNO_32124862.1:682598-682690, 2R, 682598, 682690, 3.4e-46, 183.0 # many TE match, 5+ exons, main predict is TE TEmatch: CG_NCBI_GNO_32105474.1:244820-245251, 2R, 244820, 245251, 4.6e-246, 846.0 # many TE match, main predict is NOT TE; scores, align are low TEmatch: CG_NCBI_GNO_32052943.1:3290288-3291277, 2R, 9617384, 9617345, 8.8e-11, 65.0 #.... test using blast (psi) w/ transposon lib (prot?) to find/mask TE repeats (note dpulex1.fa is repeatmasked; not Dros.spp though) /bio/bio-grid/tmpd/TransposonPSI $nb/blastall -d dmelchr4.fa -i transposon_PSI_LIB/gypsy.refSeq -R transposon_PSI_LIB/gypsy.chkp \ -m 9 -o dmelchr4.gypsy.tpsiblout \ -p psitblastn -F F -M BLOSUM45 -t -1 -e 1e-5 -v 10000 -b 10000 == 37 blast matches grep '^4' $em/dmel4c/dmel*transposon.gff | grep gypsy | wc = 11 matches; not the same as the psiblast ones though :( ** So, Transposon matches are a problem for Dmel (non-masked fa); repeat-masked Dpulex probably isnt showing much TE matching; Dmoj, other Dros species are a question; the exon-blat near/far stats don't suggest much TE match, but some repetitive genes could be in the bunch of nonTE genes; Do we want to filter and/or mask, and/or use ReAS masks (some of which may be bogus masking of true coding genes)? >> most of these are FBtr matches (2997 match = ncrna ; should have been weeded); some are predictors; enough to worry? NO> these could be mostly short matches (20bp) chr2L: 2997 FBtr; 744 NCBI; 778 DGIL; 777 RGUI; 427 BATZ; 297 TRdmel ; ** chr2R: 2219 FBtr; 3302 NCBI; 3188 DGIL; 1582 BATZ; 1425 TRdmel; Use this blat option with dmel4c.fa lower mask of transposon regions: -mask=type Mask out repeats. Alignments won't be started in masked region .. Types are lower - mask out lower cased sequence ** .repeat masker style mask file done: puma:/home/gilbertd/bio/dmel4/dmel-te-mask.out ch="4"; $kn/maskOutFa $sc/dmel2/perchr/$ch.fa dmel-te-mask.out dmel$ch-note.fa.masked #.................................................... # C. elegans, with alt transcripts parsing perl $td/tandynear.perl -debug -skip cele1.ref.dupids -nogroup \ $em/cele1/{I,V,X}*/cele1_exons.nr.blatf8 $em/cele1/{I,V,X}*/cele1_exons.nr \ > & ! cele1-tandynear-nodups.txt & # exon_fasta_ids ref=X ngroup=1; ngenes=2743; ntrans=3788; nexons=24169; naltexons=12394 # exon_fasta_ids ref=V ngroup=1; ngenes=4930; ntrans=6026; nexons=30065; naltexons=10110 # exon_fasta_ids ref=IV ngroup=1; ngenes=3210; ntrans=4505; nexons=23568; naltexons=12224 # exon_fasta_ids ref=III ngroup=1; ngenes=2578; ntrans=3951; nexons=20580; naltexons=12173 # exon_fasta_ids ref=II ngroup=1; ngenes=3401; ntrans=4675; nexons=23497; naltexons=11423 # exon_fasta_ids ref=I ngroup=1; ngenes=2785; ntrans=4092; nexons=22443; naltexons=12268 # skipped ids n=207937 Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside all 185403 157499 8430 2917 2228 13083 1246 all freq 1.000 0.054 0.019 0.014 0.083 0.008 all found 1.000 0.410 0.461 0.446 0.331 0.196 # cele1, -minalign 0.9, no dup alt-tr/exons Group Total Same Near8k Near16k Near48k Far Inside all 173240 157499 4727 1871 1356 7254 533 all freq 1.000 0.030 0.012 0.009 0.046 0.003 all found 1.000 0.715 0.701 0.716 0.582 0.411 #.................................................... # cele2, C. elegans, with ref and 2 predictor groups, overlapping transcript exon removal perl $td/tandynear.perl -debug -group \ $em/cele2/{I,V,X}*/cele1_exons.nr.blatf8 $em/cele2/{I,V,X}*/cele1_exons.nr \ > & ! cele2-tandynear.txt & # exon_fasta_ids ref=X ngroup=3; ngenes=9057; ntrans=9216; nexons=43868; naltexons=52469 # exon_fasta_ids ref=V ngroup=3; ngenes=15111; ntrans=15283; nexons=59198; naltexons=66426 # exon_fasta_ids ref=IV ngroup=3; ngenes=10568; ntrans=10737; nexons=43107; naltexons=50815 # exon_fasta_ids ref=III ngroup=3; ngenes=8390; ntrans=8573; nexons=36209; naltexons=43546 # exon_fasta_ids ref=II ngroup=3; ngenes=10673; ntrans=10847; nexons=42398; naltexons=49255 # exon_fasta_ids ref=I ngroup=3; ngenes=9258; ntrans=9427; nexons=40687; naltexons=48698 cele, Tandy exon match types per predictor group Group Total Same Near8k Near16k Near48k Far Inside WB_gene 364670 305891 18693 6427 5206 26243 2210 WB_gene freq 1.000 0.061 0.021 0.017 0.086 0.007 WB_gene found 1.000 0.452 0.489 0.532 0.414 0.263 WB_gene anyfnd 1.000 0.513 0.579 0.602 0.475 0.314 WB_gene fnd/any 1.000 0.882 0.845 0.883 0.872 0.839 Genefinder 365005 276294 17483 6830 5590 50140 8668 Genefinder freq 1.000 0.063 0.025 0.020 0.181 0.031 Genefinder found 1.000 0.481 0.518 0.536 0.326 0.504 Genefinder anyfnd 1.000 0.523 0.570 0.574 0.356 0.523 Genefinder fnd/any 1.000 0.919 0.910 0.934 0.914 0.964 twinscan 331919 263875 16511 6304 4725 37010 3494 twinscan freq 1.000 0.063 0.024 0.018 0.140 0.013 twinscan found 1.000 0.503 0.565 0.573 0.440 0.268 twinscan anyfnd 1.000 0.538 0.608 0.608 0.476 0.312 twinscan fnd/any 1.000 0.935 0.930 0.943 0.924 0.859 #.................................................... # does cele have alt-tr/exon overlaps? # yes, 46917 : could be throwing above stats off (many more Same matches) # ....... # Tandy exon match types per predictor group # Group Total Same Near Far Inside # all 145314 122807 9152 12434 921 # all freq 1.000 0.075 0.101 0.007 << freq more suggestive of tandems than before removing alt-tr exons # all found 1.000 0.375 0.316 0.194 .. i.e. due to 1/2 fewer Same exon matches gzcat $em/cele1/{I,V,X}*/cele1_exons.fa.gz | grep '^>' | perl -ne\ '($d)=m/^>(\S+)/; ($d,$x,$b,$e)=split(/[\.:-]/,$d); ($r)=m/loc=(\w+)/; print join("\t",$b,$e,$r,$d,$x),"\n";' \ | sort -k3,3 -k1,1n -k2,2nr -k4,4 -k5,5n \ | perl -ne'($b,$e,$r,$d,$x)=split; print "$d.$x\n" if($r eq $lr && $b <= $le && $e >= $lb); ($lr,$lb,$le)= ($r,$b,$e);' \ > cele1.ref.dupids #.................................................... # Apis mell. perl $td/tandynear.perl -nogroup $em/amel4/chr*/amel4_exons.nr.blatf8 $em/amel4/chr*/amel4_exons.nr Tandy exon match types per predictor group Group Total Same Near Far all 47565 46225 719 600 all freq 1.000 0.016 0.013 all found 1.000 0.234 0.323 =item R stats tn <- read.csv(stdin(),comment.char="#") # tandynear stats of exon matches to genomes (by blat, >50% align) Group,Total,Same,Near8k,Near16k,Near48k,Far,Inside Daphx_GNO,236201,177386,10748,6522,8863,29618,3064 Drosmoj_GNO,125055,112388,2910,1001,1199,7094,463 Drosere_GNO,246689,224803,4629,2007,2458,12025,767 Drosmel_GNO,254883,181388,10431,10679,24411,27446,528 Drosmel_FB,193434,175278,5085,2890,4846,4868,467 Celeg_WB,145314,122807,7041,2222,1748,10575,921 tn1 <- tn[,3:ncol(tn)] rownames(tn1) <- tn[,1] tnt <- t(tn1) class(tnt)<-"table" tndf <- as.data.frame(tnt) colnames(tndf)<-c("Dup_distance","Organism","Freq") tnx <- xtabs( Freq ~ Dup_distance + Organism, tndf) CrossTable( tnx, expected = T, dnn=c("Dup_distance","Organism"),format="SPSS",prop.t=F, digits=1 ) Cell Contents |-------------------------| | Count | | Expected Values | | Chi-square contribution | | Row Percent | | Column Percent | |-------------------------| Total Observations in Table: 1201576 | Organism Dup_distance | Daphx_GNO | Drosmoj_GNO | Drosere_GNO | Drosmel_GNO | Drosmel_FB | Celeg_WB | Row Total | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Same | 177386 | 112388 | 224803 | 181388 | 175278 | 122807 | 994050 | | 195406.4 | 103456.6 | 204083.0 | 210861.8 | 160025.7 | 120216.6 | | | 1661.8 | 771.1 | 2103.7 | 4119.8 | 1453.7 | 55.8 | | | 17.8% | 11.3% | 22.6% | 18.2% | 17.6% | 12.4% | 82.7% | | 75.1% | 89.9% | 91.1% | 71.2% | 90.6% | 84.5% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Near8k | 10748 | 2910 | 4629 | 10431 | 5085 | 7041 | 40844 | | 8029.0 | 4250.9 | 8385.5 | 8664.0 | 6575.2 | 4939.5 | | | 920.8 | 423.0 | 1682.8 | 360.4 | 337.7 | 894.1 | | | 26.3% | 7.1% | 11.3% | 25.5% | 12.4% | 17.2% | 3.4% | | 4.6% | 2.3% | 1.9% | 4.1% | 2.6% | 4.8% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Near16k | 6522 | 1001 | 2007 | 10679 | 2890 | 2222 | 25321 | | 4977.5 | 2635.3 | 5198.5 | 5371.2 | 4076.3 | 3062.2 | | | 479.3 | 1013.5 | 1959.4 | 5245.2 | 345.2 | 230.5 | | | 25.8% | 4.0% | 7.9% | 42.2% | 11.4% | 8.8% | 2.1% | | 2.8% | 0.8% | 0.8% | 4.2% | 1.5% | 1.5% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Near48k | 8863 | 1199 | 2458 | 24411 | 4846 | 1748 | 43525 | | 8556.0 | 4529.9 | 8935.9 | 9232.7 | 7006.8 | 5263.7 | | | 11.0 | 2449.3 | 4696.0 | 24952.7 | 666.4 | 2348.2 | | | 20.4% | 2.8% | 5.6% | 56.1% | 11.1% | 4.0% | 3.6% | | 3.8% | 1.0% | 1.0% | 9.6% | 2.5% | 1.2% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Far | 29618 | 7094 | 12025 | 27446 | 4868 | 10575 | 91626 | | 18011.5 | 9536.1 | 18811.2 | 19436.1 | 14750.3 | 11080.9 | | | 7479.2 | 625.4 | 2448.2 | 3301.0 | 6620.9 | 23.1 | | | 32.3% | 7.7% | 13.1% | 30.0% | 5.3% | 11.5% | 7.6% | | 12.5% | 5.7% | 4.9% | 10.8% | 2.5% | 7.3% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Inside | 3064 | 463 | 767 | 528 | 467 | 921 | 6210 | | 1220.7 | 646.3 | 1274.9 | 1317.3 | 999.7 | 751.0 | | | 2783.3 | 52.0 | 202.4 | 472.9 | 283.9 | 38.5 | | | 49.3% | 7.5% | 12.4% | 8.5% | 7.5% | 14.8% | 0.5% | | 1.3% | 0.4% | 0.3% | 0.2% | 0.2% | 0.6% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Column Total | 236201 | 125055 | 246689 | 254883 | 193434 | 145314 | 1201576 | | 19.7% | 10.4% | 20.5% | 21.2% | 16.1% | 12.1% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Statistics for All Table Factors Pearson's Chi-squared test ------------------------------------------------------------ Chi^2 = 83511.9 d.f. = 25 p = 0 Minimum expected frequency: 646.3108 #.. without Same cat; tn2 <- tn[,4:ncol(tn)] rownames(tn2) <- tn[,1] tnt2 <- t(tn2) class(tnt2)<-"table" tndf2 <- as.data.frame(tnt2) colnames(tndf2)<-c("Dup_distance","Organism","Freq") tnx2 <- xtabs( Freq ~ Dup_distance + Organism, tndf2) CrossTable( tnx2, expected = T, dnn=c("Dup_distance","Organism"),format="SPSS",prop.t=F, digits=1 ) Cell Contents |-------------------------| | Count | | Expected Values | | Chi-square contribution | | Row Percent | | Column Percent | |-------------------------| Total Observations in Table: 207526 | Organism Dup_distance | Daphx_GNO | Drosmoj_GNO | Drosere_GNO | Drosmel_GNO | Drosmel_FB | Celeg_WB | Row Total | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Near8k | 10748 | 2910 | 4629 | 10431 | 5085 | 7041 | 40844 | | 11575.6 | 2493.0 | 4307.5 | 14464.8 | 3573.4 | 4429.7 | | | 59.2 | 69.7 | 24.0 | 1124.9 | 639.5 | 1539.4 | | | 26.3% | 7.1% | 11.3% | 25.5% | 12.4% | 17.2% | 19.7% | | 18.3% | 23.0% | 21.2% | 14.2% | 28.0% | 31.3% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Near16k | 6522 | 1001 | 2007 | 10679 | 2890 | 2222 | 25321 | | 7176.2 | 1545.5 | 2670.4 | 8967.4 | 2215.3 | 2746.2 | | | 59.6 | 191.9 | 164.8 | 326.7 | 205.5 | 100.0 | | | 25.8% | 4.0% | 7.9% | 42.2% | 11.4% | 8.8% | 12.2% | | 11.1% | 7.9% | 9.2% | 14.5% | 15.9% | 9.9% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Near48k | 8863 | 1199 | 2458 | 24411 | 4846 | 1748 | 43525 | | 12335.4 | 2656.7 | 4590.2 | 15414.3 | 3807.9 | 4720.5 | | | 977.5 | 799.8 | 990.4 | 5251.0 | 283.0 | 1871.7 | | | 20.4% | 2.8% | 5.6% | 56.1% | 11.1% | 4.0% | 21.0% | | 15.1% | 9.5% | 11.2% | 33.2% | 26.7% | 7.8% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Far | 29618 | 7094 | 12025 | 27446 | 4868 | 10575 | 91626 | | 25967.7 | 5592.7 | 9663.0 | 32449.2 | 8016.2 | 9937.2 | | | 513.1 | 403.0 | 577.4 | 771.4 | 1236.4 | 40.9 | | | 32.3% | 7.7% | 13.1% | 30.0% | 5.3% | 11.5% | 44.2% | | 50.4% | 56.0% | 54.9% | 37.3% | 26.8% | 47.0% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Inside | 3064 | 463 | 767 | 528 | 467 | 921 | 6210 | | 1760.0 | 379.0 | 654.9 | 2199.3 | 543.3 | 673.5 | | | 966.2 | 18.6 | 19.2 | 1270.0 | 10.7 | 91.0 | | | 49.3% | 7.5% | 12.4% | 8.5% | 7.5% | 14.8% | 3.0% | | 5.2% | 3.7% | 3.5% | 0.7% | 2.6% | 4.1% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Column Total | 58815 | 12667 | 21886 | 73495 | 18156 | 22507 | 207526 | | 28.3% | 6.1% | 10.5% | 35.4% | 8.7% | 10.8% | | -------------|-------------|-------------|-------------|-------------|-------------|-------------|-------------| Statistics for All Table Factors Pearson's Chi-squared test ------------------------------------------------------------ Chi^2 = 20596.58 d.f. = 20 p = 0 Minimum expected frequency: 379.0468 =item fixed bug odd: not finding unpredicted exons I know are there from map views ... ** problem was we used exon-pred location (qb,qe) not genome match loc (tb,te) correct table with genome (tb,te) locations melon.% cat $em/daphd/scaffold_4/dpulex1_exons.nr.blatf8 | perl $td/tandynear.perl -fa $em/daphd/scaffold_4/dpulex1_exons.nr Tandy exon match types per predictor group Group Total Same Near Far DP_DGIL_SNO_ 35241 16328 3930 14983 DP_DGIL_SNO_ found 0.000 0.419 0.151 DP_DGIL_SNO_ anyfnd 0.000 0.465 0.158 Dappu 9402 7144 1292 966 Dappu found 0.000 0.220 0.226 << the missing Dappu predictions for tandems Dappu anyfnd 0.000 0.490 0.310 NCBI_GNO_ 12149 7840 2022 2287 NCBI_GNO_ found 0.000 0.453 0.221 NCBI_GNO_ anyfnd 0.000 0.485 0.254 melon.% cat scaffold_6680/dmoj_caf060210_exons.nr.blatf8 | perl $td/tandynear.perl -fa scaffold_6680/dmoj_caf060210_exons.nr Tandy exon match types per predictor group Group Total Same Near Far GI_BREN_NSC_ 18007 14817 450 2740 GI_BREN_NSC_ found 0.000 0.320 0.016 GI_BREN_NSC_ anyfnd 0.000 0.391 0.038 GI_DGIL_SNO_ 35983 24270 782 10931 GI_DGIL_SNO_ found 0.000 0.263 0.039 GI_DGIL_SNO_ anyfnd 0.000 0.317 0.047 GI_NCBI_GNO_ 23024 21493 703 828 GI_NCBI_GNO_ found 0.000 0.427 0.016 GI_NCBI_GNO_ anyfnd 0.000 0.444 0.024 GLEAN_ 25734 21967 697 3070 GLEAN_ found 0.000 0.374 0.038 GLEAN_ anyfnd 0.000 0.438 0.072 =cut