=item NOTE on results The main result is identification of adjacent duplicated and conserved (multi-)coding-exon regions. Tandem gene calls can, likely do, include pseudogenes and other mixed models among this. =item single exon genes are a problem Many of the tests and grouping filters here are looking at multi-exon runs, so single exon genes are not handled well. This is a problem particularly for fruitfly data, where 1-exon genes are about 50% of genome, versus 10% or less for mouse, worm, daphnia. =item Daphnia pulex tests 2007 june 14 melon:/bio/bio-grid/mb/EVidenceModeler/daphd scaffold_1 size=4193030 tandy HSP=739 ; match=161 ; region=90 scaffold_2 size=3740169 tandy HSP=1068 ; match=268 ; region=153 scaffold_3 size=3777634 tandy HSP=1281 ; match=304 ; region=153 scaffold_4 size=3075709 tandy HSP=1785 ; match=421 ; region=191 scaffold_5 size=2511979 tandy HSP=1028 ; match=172 ; region=76 scaffold_6 size=2406117 tandy HSP=1475 ; match=367 ; region=173 scaffold_7 size=2324446 tandy HSP=960 ; match=227 ; region=101 scaffold_8 size=2335496 tandy HSP=695 ; match=157 ; region=82 scaffold_9 size=2251199 tandy HSP=1259 ; match=353 ; region=201 26 MB total Gene counts cat scaffold_*/dpulex1_predict.gff | grep mRNA | perl -ne'($r,$s,$t)=split; print "$s\n" if($t eq "mRNA");' | sort | uniq -c 11196 DGIL_SNO << this is 2x data stutter; 5598 is right 3315 JGI 4332 NCBI_GNO 2381 twinscan ?? drop these use 4332 as median gene count # genebest for Dpulex scaffold 1-9 # gene bestids scores/method (num is all ids x all gene matches, ~ 2/match) predictors num score lost tandem near equal DP_DGIL_SNO_: 3945 65.92 6.77 0.70 0.23 0.33 Dappu_FM_: 1819 63.80 2.47 0.63 0.24 0.31 NCBI_GNO_: 3110 64.18 2.93 0.88 0.26 0.30 # geneinfo for Dpulex scaffold 1-9 # gene info, all-methods scores/method (count is all gene matches) predictor count eqfar eqnear eqsame fullgns gnids maxgns methods nexons tandkno tandnew tandems DP_DGIL_SNO_ 2049 0.89 1.41 3.77 1.72 4.08 2.12 1.87 4.67 374 98 0.50 Dappu_FM_ 911 1.23 2.25 6.09 3.36 7.12 3.05 2.71 7.08 284 44 0.76 NCBI_GNO_ 1322 1.04 1.90 4.96 2.51 5.66 2.60 2.37 5.92 332 81 0.68 all 2430 0.87 1.26 3.32 1.50 3.65 1.97 1.76 4.23 387 114 0.47 # ............................................................. Using 501 for all known/new tandems, and 4332 gene predictions for these 9 scaffolds gives tandem rate of ~ 12% (is this calc right? known tandems are included in base count of 4000, but not 'new' tandem models) #----------- Dpulex, daphe : all scaffolds, tandy v6g ---------------- ** scaffolds excluding <10kb + bacteria, n=1207 ** NOTE: For comparing to Cele, Dmel with only 1 gene class (no multi predictors), ** rerun with -igclass to keep only NCBI_GNO Prediction counts # cat scaffold_*/dpulex1_predict.gff | grep mRNA | perl -ne'($r,$s,$t)=split; print "$s\n" if($t eq "mRNA");' | sort | uniq -c 37818 DGIL_SNO 17351 JGI (excluding SNAP in DGIL) 27084 NCBI_GNO .. use NCBI_GNO as best count & predictor ## gene counts in 1st 100 scaffolds: 9 scaffolds: 25227 DGIL_SNO 5598 DGIL_SNO 12792 JGI 3315 JGI 18843 NCBI_GNO 4332 NCBI_GNO ** part/most of diff in match count for dpulex,dmoj vs dmel,cele is multiple predictors; test w/ only NCBI_GNO # all species, using 4 predictors for dpulex,dmoj (exons_tandy6g.gff) #suspect# dpulex: 25915 / 10459 match/region = 2.5 #suspect# dmoj : 23981 / 11324 match/region = 2.1 dmel : 2209 / 969 match/region = 2.3 // some of these are ncRNA (tRNA+), very repeated. cele : 7754 / 3215 match/region = 2.4 # .. dpulex, dmoj using only NCBI_GNO predictions (exons_tandy6i.gff) dmoj : 6922 / 3208 match/region = 2.1 dpulex: 13010 / 5831 match/region = 2.2 ** use absolute count of tandems as measure among species? # cat */*_exons_tandy6i.gff | grep -v tandyweak | grep bestid | grep -vc eqnear:0 dmel (ref): eqnear: 1352/2139 ; eqfar: 562/2139 ; eqsame: 1628/2139 dmoj (GNO): eqnear: 2677/6042 ; eqfar: 2039/6042 ; eqsame: 3587/6042 cele (ref): eqnear: 3392/7199 ; eqfar: 2816/7199 ; eqsame: 4638/7199 dpulex (GNO): eqnear: 4708/11945 ; eqfar: 4882/11945 ; eqsame: 7254/11945 # count of rich regions: region with 3+ genes, each with large size (measure against gene size?) # gene clusters (n>=3), near/far/same counts dmel (ref): total: 215 eqnear: 205 eqfar: 85 eqsame: 212 dmoj (GNO): total: 487 eqnear: 340 eqfar: 279 eqsame: 393 cele (ref): total: 846 eqnear: 721 eqfar: 434 eqsame: 788 dplx (GNO): total: 1137 eqnear: 857 eqfar: 723 eqsame: 1039 #.......... # gene cluster counts /bio/bio-grid/mb/EVidenceModeler/daphe melon.% cat */*_exons_tandy6i.gff | grep bestids= | env nc=2 perl -n $td/geneclust.perl # gene clusters (n>=2), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 2840 eqnear: 1934 eqfar: 1454 eqsame: 2529 duponly: 1345 average genes/region: eqnear: 1.8 eqfar: 2.0 eqsame: 2.0 duponly: 1.4 # gene clusters (n>=3), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 919 eqnear: 747 eqfar: 584 eqsame: 861 duponly: 509 average genes/region: eqnear: 2.6 eqfar: 2.7 eqsame: 2.8 duponly: 1.6 # gene clusters (n>=4), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 417 eqnear: 359 eqfar: 304 eqsame: 390 duponly: 241 average genes/region: eqnear: 3.3 eqfar: 3.2 eqsame: 3.6 duponly: 1.8 /bio/bio-grid/mb/EVidenceModeler/cele1 melon.% cat */*_exons_tandy6g.gff | grep bestids= | env nc=2 perl -n $td/geneclust.perl # gene clusters (n>=2), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 1662 eqnear: 1310 eqfar: 739 eqsame: 1508 duponly: 588 average genes/region: eqnear: 2.2 eqfar: 2.0 eqsame: 2.4 duponly: 1.4 # gene clusters (n>=3), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 680 eqnear: 612 eqfar: 344 eqsame: 658 duponly: 219 average genes/region: eqnear: 2.9 eqfar: 2.4 eqsame: 3.4 duponly: 1.6 # gene clusters (n>=4), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 290 eqnear: 272 eqfar: 175 eqsame: 281 duponly: 116 average genes/region: eqnear: 3.9 eqfar: 3.0 eqsame: 4.3 duponly: 1.8 /bio/bio-grid/mb/EVidenceModeler/dmoj2 melon.% cat */*_exons_tandy6i.gff | grep bestids= | env nc=2 perl -n $td/geneclust.perl # gene clusters (n>=2), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 1142 eqnear: 874 eqfar: 458 eqsame: 981 duponly: 319 average genes/region: eqnear: 1.7 eqfar: 2.3 eqsame: 2.2 duponly: 1.6 # gene clusters (n>=3), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 338 eqnear: 247 eqfar: 216 eqsame: 283 duponly: 138 average genes/region: eqnear: 2.8 eqfar: 3.0 eqsame: 3.1 duponly: 2.0 # gene clusters (n>=4), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 171 eqnear: 124 eqfar: 123 eqsame: 146 duponly: 80 average genes/region: eqnear: 3.6 eqfar: 3.7 eqsame: 3.7 duponly: 2.1 /bio/bio-grid/mb/EVidenceModeler/dmel4 # gene clusters (n>=2), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 474 eqnear: 436 eqfar: 141 eqsame: 451 duponly: 78 average genes/region: eqnear: 2.2 eqfar: 2.7 eqsame: 2.4 duponly: 1.7 # gene clusters (n>=3), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 168 eqnear: 162 eqfar: 71 eqsame: 167 duponly: 36 average genes/region: eqnear: 3.2 eqfar: 3.7 eqsame: 3.4 duponly: 2.0 # gene clusters (n>=4), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 73 eqnear: 73 eqfar: 38 eqsame: 73 duponly: 24 average genes/region: eqnear: 4.2 eqfar: 4.8 eqsame: 4.2 duponly: 2.4 # gene info & cluster counts, alternate predictors /bio/bio-grid/mb/EVidenceModeler/daphe JGI FM5 models # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgn maxgn nexon align tand tandat tandnew duplat duplnew Dappu 7014 0.84 0.74 1.60 0.55 1.66 3.26 72.04 0.22 1182 340 1560 1193 Using JGI Dappu: 1522 for known/new tandems, and 17351 JGI genes, gives tandem rate of 8.7% (but 1/2 total of NCBI_GNO) melon.% cat sca*/*_exons_tandy6i_jgi.gff | grep bestids= | env nc=2 perl -n $td/geneclust.perl # gene clusters (n>=2), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 1731 eqnear: 1122 eqfar: 881 eqsame: 1437 duponly: 1018 average genes/region: eqnear: 1.6 eqfar: 1.9 eqsame: 1.7 duponly: 1.5 # gene clusters (n>=3), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 471 eqnear: 342 eqfar: 309 eqsame: 394 duponly: 322 average genes/region: eqnear: 2.3 eqfar: 2.5 eqsame: 2.3 duponly: 1.9 # gene clusters (n>=4), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 180 eqnear: 141 eqfar: 132 eqsame: 150 duponly: 133 average genes/region: eqnear: 3.0 eqfar: 3.0 eqsame: 2.7 duponly: 2.3 /bio/bio-grid/mb/EVidenceModeler/dmoj2 BRENT_NSCAN models Note: BREN_NSC produced slightly more total predictions than NCBI_GNO here, but only 1/2 tandem dupls # gene info, all-methods predictor count eqfar eqnear eqsame fullgn maxgn nexon align tand tandat tandnew duplat duplnew GI_BREN_NSC_ 4539 0.86 0.42 0.96 0.82 2.52 2.51 59.12 0.18 378 85 625 517 Using BREN_NSC: 463 for known/new tandems, and 17340 NSC genes (12 major scaffs) gives tandem rate of 2.6% melon.% cat sca*/*_exons_tandy6i_nsc.gff | grep bestids= | env nc=2 perl -n $td/geneclust.perl # gene clusters (n>=2), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 705 eqnear: 366 eqfar: 463 eqsame: 546 duponly: 329 average genes/region: eqnear: 2.0 eqfar: 1.9 eqsame: 2.1 duponly: 1.4 # gene clusters (n>=3), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 254 eqnear: 168 eqfar: 171 eqsame: 221 duponly: 123 average genes/region: eqnear: 2.7 eqfar: 2.5 eqsame: 2.8 duponly: 1.7 # gene clusters (n>=4), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 127 eqnear: 91 eqfar: 89 eqsame: 115 duponly: 70 average genes/region: eqnear: 3.3 eqfar: 3.1 eqsame: 3.3 duponly: 1.8 # Honey bee / Apis mellifera, using NCBI Gnomon models /bio/bio-grid/mb/EVidenceModeler/amel4 cat chr*/amel4_exons_tandy6g.gff | grep bestids= | env nc=2 perl -n $td/geneclust.perl # gene clusters (n>=2), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 132 eqnear: 98 eqfar: 42 eqsame: 126 duponly: 37 average genes/region: eqnear: 2.2 eqfar: 1.9 eqsame: 2.4 duponly: 1.6 # gene clusters (n>=3), near/far/same counts total regions: 54 eqnear: 48 eqfar: 24 eqsame: 53 duponly: 14 average genes/region: eqnear: 3.0 eqfar: 2.2 eqsame: 3.4 duponly: 2.4 # gene clusters (n>=4), near/far/same counts total regions: 26 eqnear: 24 eqfar: 16 eqsame: 25 duponly: 10 average genes/region: eqnear: 4.0 eqfar: 2.4 eqsame: 4.1 duponly: 2.8 #...................... # cat scaffold_*/*_exons_tandy6g.gff | grep bestids= | perl -n $td/geneinfo.perl # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgn maxgn nexon align tand tandat tandnew duplat duplnew DP_DGIL_SNO_ 19578 1.20 1.06 2.99 1.24 2.50 4.47 73.41 0.41 4451 561 6639 2248 Dappu 11979 1.28 1.31 3.76 1.72 2.85 5.26 70.25 0.48 3577 225 5035 944 NCBI_GNO_ 16192 1.20 1.17 3.34 1.45 2.65 4.81 71.24 0.44 4315 333 6206 1462 all 23642 1.08 0.95 2.68 1.11 2.30 4.11 73.33 0.37 4886 645 7237 2677 # ............................................................. Using NCBI_GNO: 4648 for all known/new tandems, and 27084 gene predictions for scaffolds gives tandem rate of 17% Using NCBI_GNO: 7668 for known/new duplicates, .., duplicate (same chr) rate of 28% Using all: 5531 for all known/new tandems, and 27084 gene predictions for scaffolds gives tandem rate of 20% Using all: 9914 for known/new duplicates, .., duplicate (same chr) rate of 37% #----------- Dpulex, daphe, NCBI_GNO only run filter : all scaffolds, tandy v6g ---------------- # find about 1/2 of dupl. when using all predictors; probably algo error for multi-predictors # cat scaffold_*/dpulex1_exons_tandy6i.gff | grep bestids= | perl -n $td/geneinfo.perl # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgn maxgn nexon align tand tandat tandnew duplat duplnew NCBI_GNO_ 11945 0.89 0.83 1.83 0.61 2.15 3.34 72.03 0.30 2332 505 3228 1689 all 11945 0.89 0.83 1.83 0.61 2.15 3.34 72.03 0.30 2332 505 3228 1689 # ............................................................. Using NCBI_GNO: 2837 for known/new tandems, and 27084 predictions, gives tandem rate of 10.4% Using NCBI_GNO: 4917 for known/new duplicates, .., duplicate (same chr) rate of 18% # ............................................................. echo "# genebest" cat scaffold_*/*_exons_tandy.gff | grep bestids= | perl -n $td/genebest.perl echo "# geneinfo" cat scaffold_*/*_exons_tandy.gff | grep bestids= | perl -n $td/geneinfo.perl echo "# ............................................................." melon.% grep -c bestids scaffold_*/*_exons_tandy6g.gff | sed -e's/:/ /' | sort -k2,2nr | more scaffold_17/dpulex1_exons_tandy6g.gff 1397 ** lots of repeats on this scaffold ** 761 vs 636 nonrep scaffold_4/dpulex1_exons_tandy6g.gff 646 scaffold_6/dpulex1_exons_tandy6g.gff 620 scaffold_18/dpulex1_exons_tandy6g.gff 602 scaffold_36/dpulex1_exons_tandy6g.gff 575 scaffold_10/dpulex1_exons_tandy6g.gff 548 scaffold_16/dpulex1_exons_tandy6g.gff 512 scaffold_12/dpulex1_exons_tandy6g.gff 489 scaffold_9/dpulex1_exons_tandy6g.gff 462 scaffold_2/dpulex1_exons_tandy6g.gff 416 scaffold_3/dpulex1_exons_tandy6g.gff 388 scaffold_29/dpulex1_exons_tandy6g.gff 365 scaffold_13/dpulex1_exons_tandy6g.gff 335 scaffold_7/dpulex1_exons_tandy6g.gff 325 scaffold_14/dpulex1_exons_tandy6g.gff 291 scaffold_1/dpulex1_exons_tandy6g.gff 286 scaffold_27/dpulex1_exons_tandy6g.gff 268 scaffold_15/dpulex1_exons_tandy6g.gff 266 scaffold_40/dpulex1_exons_tandy6g.gff 265 scaffold_70/dpulex1_exons_tandy6g.gff 258 scaffold_25/dpulex1_exons_tandy6g.gff 249 scaffold_8/dpulex1_exons_tandy6g.gff 249 scaffold_24/dpulex1_exons_tandy6g.gff 237 scaffold_72/dpulex1_exons_tandy6g.gff 232 scaffold_60/dpulex1_exons_tandy6g.gff 228 scaffold_5/dpulex1_exons_tandy6g.gff 221 melon.% grep -c region scaffold_*/*_exons_tandy6g.gff | sed -e's/:/ /' | sort -k2,2nr | more scaffold_6/dpulex1_exons_tandy6g.gff 231 scaffold_10/dpulex1_exons_tandy6g.gff 217 scaffold_4/dpulex1_exons_tandy6g.gff 206 scaffold_9/dpulex1_exons_tandy6g.gff 199 scaffold_2/dpulex1_exons_tandy6g.gff 196 scaffold_18/dpulex1_exons_tandy6g.gff 195 scaffold_12/dpulex1_exons_tandy6g.gff 194 scaffold_3/dpulex1_exons_tandy6g.gff 187 scaffold_17/dpulex1_exons_tandy6g.gff 186 scaffold_36/dpulex1_exons_tandy6g.gff 164 scaffold_16/dpulex1_exons_tandy6g.gff 162 scaffold_27/dpulex1_exons_tandy6g.gff 159 scaffold_1/dpulex1_exons_tandy6g.gff 139 scaffold_13/dpulex1_exons_tandy6g.gff 137 scaffold_54/dpulex1_exons_tandy6g.gff 131 scaffold_7/dpulex1_exons_tandy6g.gff 131 scaffold_22/dpulex1_exons_tandy6g.gff 130 scaffold_26/dpulex1_exons_tandy6g.gff 120 scaffold_14/dpulex1_exons_tandy6g.gff 112 scaffold_8/dpulex1_exons_tandy6g.gff 109 scaffold_34/dpulex1_exons_tandy6g.gff 106 scaffold_5/dpulex1_exons_tandy6g.gff 101 scaffold_67/dpulex1_exons_tandy6g.gff 101 =item Tandem gene Genome biology http://www.wormbook.org/chapters/www_geneduplication/geneduplication.html Gene duplications are very widespread in C. elegans and appear to arise more frequently than in either Drosophila or yeast. ... The total number of genes in the paranome is around 6,000, or 32% of the genome, where the paranome is defined as the set of proteins that have one or more paralogues, ... Local clusters of duplicated genes are easily identified and there are 402 such clusters (of 3+ genes) -- lots of good detail here ............... =item C. elegans See gene-struct stats: Celegans is closer to Daphnia than Ffly, Mouse, .. Need another organism with multi=exon genes similar to Daphnia 2-exon average in flies has problems w/ some of the classifying that relies on >1 exon/gene Use this genome data: microbe% zgrep mRNA $sc/cele1/celegans-slim-WS167.gff.gz | perl -ne'($r,$s,$t,@x)=split; print "$s\n" if($t eq "mRNA");' | sort | uniq -c 27049 Coding_transcript << these mRNA 21182 Genefinder 221 Transposon_CDS 8641 history 21681 twinscan ftp://ftp.wormbase.org/ wormbase/genomes/elegans/sequences/dna/ wormbase/genomes/elegans/genome_feature_tables/GFF3/ zcat elegansWS167.gff3.gz | egrep -v '_match|RNAi_reagent|SAGE_tag|intron|_repeat' > celegansWS167slim.gff Sizes (note there are ~ 2 mRNA/gene) III 13783317 genes=2664 ; tandy HSP=1804 ; match=462 ; region=210 II 15279313 genes=3484 ; tandy HSP=4473 ; match=1253 ; region=534 I 15072419 genes=2866 ; tandy HSP=2248 ; match=569 ; region=255 IV 17493785 genes=3282 ; tandy HSP=3761 ; match=1018 ; region=473 V 20922231 genes=4997 ; tandy HSP=7970 ; match=2363 ; region=1053 X 17718851 genes=2784 ; tandy HSP=1911 ; match=511 ; region=222 MtDNA 13794 totals: 20077 genes; tandy : 22167 HSP ; 6176 match ; 2747 region cele1, gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgns gnids maxgns methods nexons tandkno tandnew tandems I: all 569 0.29 1.39 3.28 0.61 1.66 1.31 1.31 3.95 95 47 0.55 II: all 1253 0.27 1.43 2.71 0.63 1.63 1.33 1.26 3.57 218 148 0.60 III: all 462 0.29 1.39 3.23 0.64 1.60 1.34 1.23 3.90 85 31 0.60 IV: all 1018 0.35 1.31 2.99 0.78 1.69 1.33 1.28 3.69 163 103 0.57 V: all 2363 0.34 1.20 2.61 0.60 1.44 1.26 1.17 3.37 356 258 0.54 X: all 511 0.14 0.99 3.12 0.52 1.42 1.20 1.18 3.74 54 41 0.46 ............ total: all 6176 0.30 1.28 2.84 0.64 1.55 1.29 1.23 3.59 971 628 0.55 ............ Using 1599 for known/new tandems among 20077 genes (6 chrs) gives tandem rate of ~ 8% # Using 6176 duplicate match regions among 20077 genes, gives general dupl. rate of 30% # using match regions may be spurious# cele1, tandy v6g, gene info, all-methods predictor count eqfar eqnear eqsame fullgn loqual maxgn nexon align tand tandat tandnew duplat duplnew total: all 7754 0.85 1.13 2.45 0.66 0.07 1.81 3.99 64.97 0.51 1817 238 2358 718 ............ Using 2055 for known/new tandems, and 20077 genes (6 chrs) gives tandem rate of 10.2% Using 3076 for known/new duplicates, .., duplicate (same chr) rate of 15.3% # Using 7754 match regions, gives general dupl. rate of 38% # using match regions may be spurious# cat {I,II,III,IV,V,X}/*_exons_tandy6g.gff | grep bestids= | perl -n $td/geneinfo.perl | grep -v '_' =item Apis mellifera * NCBI Genbank annotation (Gnomon), release 4 genome ignoring uncategorized scaffolds (~ 2000 genes) : 6908 genes melon.% cat chrLG*/amel4_genes.gff | grep -c mRNA = 6908 # melon.% cat chrLG*/*exons_tandy6g.gff | grep bestid | perl -n $td/geneinfo.perl # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgn maxgn nexon align tand tandat tandnew duplat duplnew LOC 673 0.44 0.96 2.77 0.36 1.38 4.98 60.90 0.30 110 19 139 44 all 704 0.52 0.99 2.81 0.37 1.42 5.02 61.44 0.33 119 19 150 46 ............ Using 138 for known/new tandems, and 6908 genes gives tandem rate of 2 % Using 196 for known/new duplicates, .., duplicate (same chr) rate of 2.8 % # melon.% cat chrLG*/*exons_tandy6g.gff | grep bestid | perl -n $td/geneclust.perl # gene clusters (n>=2), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 132 eqnear: 98 eqfar: 42 eqsame: 126 duponly: 37 average genes/region: eqnear: 2.2 eqfar: 1.9 eqsame: 2.4 duponly: 1.6 # gene clusters (n>=3), near/far/same counts # criteria: fullgenes>0 or pctalign>33 and !overlap and !lowquality total regions: 54 eqnear: 48 eqfar: 24 eqsame: 53 duponly: 14 average genes/region: eqnear: 3.0 eqfar: 2.2 eqsame: 3.4 duponly: 2.4 =item Drosophila melanogaster set dpid=dmel4 * Redo this adding Dmel caf1 predictions (NCBI_GNO, DGIL_SNO, BATZ, RGUID, ..) cp -p $sc/dmel3/dmel_r430.fa.gz ${dpid}.fa.gz ; gunzip ${dpid}.fa.gz echo '##gff-version 3' > ${dpid}_predict.gff # Release 4.3 GFF gene model like this, use 'exon' not CDS; cant recode as CDS. gzcat $sc/dmel3/gff/dmel-*-r4.3.0.gff.gz | grep FlyBase | perl -ne \ '($r,$s,$t)=split"\t"; print if($t =~ /gene|mRNA|exon|CDS|protein|chromosome_arm/ && $s eq "FlyBase"); ' \ >> ${dpid}_predict.gff Sizes (note there are ~ 1.5 mRNA/gene) 2L 22407834 genes=2742 ; tandy HSP=804 ; match=264 ; region=100 2R 20766785 genes=3009 ; tandy HSP=781 ; match=316 ; region=121 3L 23771897 genes=2787 ; tandy HSP=371 ; match=149 ; region=65 3R 27905053 genes=3554 ; tandy HSP=891 ; match=314 ; region=143 4 1281640 genes=95 ; tandy HSP=66 ; match=16 ; region=7 X 22224390 genes=2313 ; tandy HSP=624 ; match=198 ; region=79 14500 genes total dmel4, gene info, all-methods predictor count eqfar eqnear eqsame fullgns gnids maxgns methods nexons tandkno tandnew tandems chr2L: all 264 0.58 1.36 2.55 5.48 6.03 6.03 1.00 3.05 53 24 1.16 chr2R: all 316 0.51 1.21 1.99 17.85 18.25 18.25 1.00 2.47 60 26 10.54 chr3L: all 149 0.34 1.13 1.83 1.50 2.28 2.28 1.00 2.49 20 22 0.98 chr3R: all 314 0.20 1.14 2.22 1.19 1.79 1.79 1.00 2.84 48 26 0.72 chrX : all 198 0.55 1.72 2.37 1.31 2.12 2.12 1.00 3.15 45 24 0.82 chr4 : all 16 1.44 1.94 3.31 0.38 1.94 1.94 1.00 4.12 2 0 0.69 ............ total: all 1257 0.45 1.30 2.22 6.33 6.93 6.93 1.00 2.81 228 122 3.33 ............ Using 350 for known/new tandems, and 14500 genes (6 chrs) gives tandem rate of ~ 2.4% dmel4, tandy v6g, gene info, all-methods predictor count eqfar eqnear eqsame fullgn loqual maxgn nexon align tand tandat tandnew duplat duplnew total: all 2209 0.45 1.23 1.90 4.35 0.03 5.29 2.97 70.16 1.76 824 44 860 111 ............ Using 868 for known/new tandems, and 14500 genes (6 chrs) gives tandem rate of 6% Using 971 for known/new duplicates, .., duplicate (same chr) rate of 6.6% cat {2L,2R,3L,3R,4,X}/*_exons_tandy6g.gff | grep bestids= | perl -n $td/geneinfo.perl | grep -v '_' =item Drosophila moj melon:/bio/bio-grid/mb/EVidenceModeler/dmoj scaffold_6541/dmoj_caf060210.fa scaf-size=2543558; Sizes: scaffold_6308 size=3356042 scaffold_6328 size=4453435 scaffold_6359 size=4525533 scaffold_6473 size=16943266 tandy HSP=2371 ; match=629 ; region=301 scaffold_6482 size=2735782 scaffold_6496 size=26866924 scaffold_6498 size=3408170 scaffold_6500 size=32352404 scaffold_6540 size=34148556 scaffold_6541 size=2543558 tandy HSP=579 ; match=212 ; region=115 scaffold_6654 size=2564135 scaffold_6680 size=24764193 tandy HSP=4477 ; match=1112 ; region=498 ** scaffold_6680 takes ~12 hrs for full tandy run (blat in ~8 hr?) #...... REDONE with 4 predictors, tandy v6g ........... melon.% cat scaffold_*/dmoj_caf060210_predict.gff | egrep 'gene|mRNA' | perl -ne'($r,$s,$t)=split; print "$s\n" if($t eq "mR NA" || ($t eq "gene" && $s eq "BREN_NSC"));' | sort | uniq -c # Main scaffolds # All scaffolds 17340 BREN_NSC 19052 BREN_NSC gene 33831 DGIL_SNO 39387 DGIL_SNO mRNA 17723 GLEAN 22451 GLEAN mRNA 14532 NCBI_GNO 17950 NCBI_GNO mRNA .. use 17723 as gene count # cat scaffold_*/*_exons_tandy6g.gff | grep bestids= | perl -n $td/geneinfo.perl dmoj2, tandy v6g, gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgn maxgn nexon align tand tandat tandnew duplat duplnew GI_BREN_NSC_ 7215 0.74 0.93 2.69 3.28 3.55 3.84 70.33 0.31 1895 89 2368 637 GI_DGIL_SNO_ 15912 0.86 0.52 1.61 1.82 2.66 2.82 70.37 0.17 2251 120 3196 2099 <* outlier GI_NCBI_GNO_ 8641 0.50 0.99 2.54 2.67 2.79 3.60 68.65 0.29 2484 71 2890 396 GLEAN_ 9175 0.59 0.91 2.43 3.09 3.50 3.59 66.20 0.29 2356 85 2864 573 all 18493 0.80 0.52 1.51 1.65 2.47 2.69 70.55 0.17 2687 145 3676 2233 # ............................................................. Using all 2832 for known/new tandems, and 17723 genes (12 major scaffs) gives tandem rate of 16% Using all 5909 for known/new duplicates, .., duplicate (same chr) rate of 33% Using NCBI_GNO: 2555 for known/new tandems, and 17723 genes (12 major scaffs) gives tandem rate of 14.5% Using NCBI_GNO: 3286 for known/new duplicates, .., duplicate (same chr) rate of 18.5% # .. Removing all but NCBI_GNO genes (Filtered at runtime) for comparison # cat scaffold_*/*_exons_tandy6i.gff | grep bestids= | perl -n $td/geneinfo.perl dmoj2, tandy v6g, gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgn maxgn nexon align tand tandat tandnew duplat duplnew GI_NCBI_GNO_ 6042 0.58 0.90 1.38 0.59 2.05 2.61 67.57 0.21 1839 95 2018 425 all 6042 0.58 0.90 1.38 0.59 2.05 2.61 67.57 0.21 1839 95 2018 425 # ............................................................. Using NCBI_GNO: 1934 for known/new tandems, and 14532 genes (12 major scaffs) gives tandem rate of 13% Using NCBI_GNO: 2443 for known/new duplicates, .., duplicate (same chr) rate of 17% #...................................................... *** REDO THIS WITH SUBSET OF PREDICTORS; now with >9 have too many to be sure what is going on with calcs, and if they are comparable (also slow) try NCBI_GNO, GLEAN, BREN_NSC, EISE_CGW, DGIL_SNO Gene counts (3 scaffs): gzgrep mRNA dmoj_caf060210_predict.gff.gz | egrep 'scaffold_6473|scaffold_6541|scaffold_6680' | \ perl -ne'($r,$s,$t)=split; print "$s\n" if($t eq "mRNA");' | sort | uniq -c 9737 DGIL_SNO 8146 EISE_CEX * multi-transcripts/gene 10313 EISE_CGW * multi-transcripts/gene 4805 GLEAN 4108 GLEANR 3826 NCBI_GNO 3862 OXFD_GPX 3441 RGUI_GID .... == ~ 4000 genes # Main scaffolds 13271 BATZ_CNA start_codon 17340 BREN_NSC gene 33831 DGIL_SNO mRNA 33013 EISE_CEX mRNA * multi-transcripts/gene 40953 EISE_CGW mRNA * multi-transcripts/gene 17723 GLEAN mRNA 15468 GLEANR mRNA 14532 NCBI_GNO mRNA 15521 OXFD_GPX mRNA 12374 RGUI_GID mRNA 44 MB total for 3 scaffs; @ st = 24764193 + 16943266 + 2543558 melon.% ll scaffold_6*/dmoj_caf060210_exons_tandy.gff 1547179 Jun 14 09:34 scaffold_6473/dmoj_caf060210_exons_tandy.gff 1098772 Jun 13 22:52 scaffold_6541/dmoj_caf060210_exons_tandy.gff 3400935 Jun 14 11:18 scaffold_6680/dmoj_caf060210_exons_tandy.gff melon.% grep -c bestids= scaffold_6*/dmoj_caf060210_exons_tandy.gff scaffold_6473/dmoj_caf060210_exons_tandy.gff:629 scaffold_6541/dmoj_caf060210_exons_tandy.gff:212 melon.% grep -c region scaffold_6*/dmoj_caf060210_exons_tandy.gff scaffold_6473/dmoj_caf060210_exons_tandy.gff:301 scaffold_6541/dmoj_caf060210_exons_tandy.gff:115 melon.% grep -c HSP scaffold_6*/dmoj_caf060210_exons_tandy.gff scaffold_6473/dmoj_caf060210_exons_tandy.gff:2371 scaffold_6541/dmoj_caf060210_exons_tandy.gff:579 # genebest for Dros.moj scaffold 6473 6541 6680 # gene bestids scores/method predictors num score lost tandem near equal GI_BATZ_CNA_: 1013 65.70 1.98 0.21 0.07 0.62 GI_BREN_NSC_: 1280 63.87 2.14 0.24 0.09 0.57 GI_DGIL_SNO_: 2666 68.56 1.34 0.43 0.09 0.51 GI_EISE_CEX_: 2256 75.30 1.76 1.07 0.18 0.71 * high tandem score here probably due to GI_EISE_CGW_: 2676 76.71 1.32 2.05 0.20 0.72 * alt transcripts of same gene (labeled as diff gene) GI_NCBI_GNO_: 1302 67.12 3.09 0.42 0.15 0.61 GI_PACH_GMP_: 1165 73.96 3.74 0.29 0.10 0.84 GI_RGUI_GID_: 1795 61.31 3.11 0.35 0.09 0.50 GLEAN_ : 1977 63.73 1.50 0.45 0.11 0.49 TRdmoj_: 1481 72.42 3.12 0.92 0.23 0.68 * alt transcripts/gene also here dmoj_GLEANR_: 1825 63.36 1.88 0.47 0.12 0.48 ** need to mark or filter out alt transcript of same gene, by overlapping exons ? # geneinfo for Dros.moj scaffold 6473 6541 6680 # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgns gnids maxgns methods nexons tandkno tandnew tandems GI_BATZ_CNA_ 635 0.60 2.66 7.14 10.66 22.21 4.75 9.79 7.39 234 9 0.78 GI_BREN_NSC_ 732 0.57 2.51 6.62 9.94 20.83 4.58 9.34 6.87 265 9 0.84 GI_DGIL_SNO_ 1400 0.42 1.41 4.40 5.65 12.11 3.06 5.59 4.69 274 19 0.49 GI_EISE_CEX_ 778 0.24 2.19 5.58 7.78 16.01 3.78 8.46 5.92 238 15 0.85 GI_EISE_CGW_ 685 0.16 2.56 6.36 9.32 18.71 4.28 9.77 6.63 245 17 1.09 GI_NCBI_GNO_ 831 0.56 2.25 5.87 9.29 19.49 4.47 8.74 6.16 265 15 0.83 GI_PACH_GMP_ 635 0.18 2.40 6.52 8.89 18.64 4.09 9.92 6.76 211 12 0.85 GI_RGUI_GID_ 993 0.51 1.99 5.34 8.42 17.80 4.26 7.93 5.62 278 14 0.80 GLEAN_ 928 0.51 2.16 5.68 9.23 19.37 4.54 8.67 5.94 285 18 0.88 TRdmoj_ 775 0.32 2.31 5.85 8.62 17.55 4.15 8.85 6.16 246 17 0.94 all 1953 0.41 1.12 3.47 4.58 9.95 2.78 4.76 3.80 302 27 0.47 dmoj_GLEANR_ 893 0.53 2.24 5.73 9.50 19.93 4.66 8.85 6.01 284 18 0.91 # ............................................................. Using 329 for all known/new tandems, and ~4000 gene predictions for these 3 scaffolds gives tandem rate of ~ 8% tandy 5q, subset predictors, most of genome (-scaffold_6540), no-repeatskip: melon.% cat scaffold_*/*_exons_tandy.gff | grep bestids= | perl -n $td/geneinfo.perl # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgn maxgn nexon tand tandat tandnew duplat duplnew GI_BREN_NSC_ 2094 0.78 1.34 3.20 3.43 2.93 3.73 0.56 465 56 514 98 GI_DGIL_SNO_ 3827 0.58 0.85 2.67 2.32 2.40 3.13 0.38 518 83 581 266 GI_NCBI_GNO_ 1967 0.74 1.47 3.45 3.66 3.11 3.88 0.70 471 59 518 96 GLEAN_ 2247 0.74 1.36 3.27 3.69 3.37 3.70 0.65 490 63 546 111 all 4819 0.57 0.75 2.27 2.01 2.16 2.77 0.35 539 107 606 322 #........ per scaf ............ melon.% grep bestids= scaffold_6541/dmoj_caf060210_exons_tandy.gff | perl -n $td/genebest.perl # gene bestids scores/method predictors num score lost tandem near equal GI_BATZ_CNA_: 261 43.48 1.20 0.28 0.02 0.09 GI_BREN_NSC_: 353 42.81 1.84 0.25 0.02 0.09 GI_DGIL_SNO_: 826 53.70 0.69 0.79 0.04 0.09 GI_EISE_CEX_: 2 64.00 0.00 0.50 0.50 0.50 GI_EISE_CGW_: 4 50.75 1.00 0.25 0.25 0.25 GI_NCBI_GNO_: 242 50.38 2.36 0.65 0.10 0.11 GI_PACH_GMP_: 4 63.50 0.00 0.25 0.25 0.25 GI_RGUI_GID_: 590 47.85 1.29 0.61 0.04 0.07 GLEAN_ : 683 49.67 0.66 0.66 0.04 0.08 TRdmoj_ : 54 71.93 0.31 0.46 0.09 0.24 dmoj_GLEANR_: 624 49.42 1.04 0.70 0.04 0.07 melon.% grep bestids= scaffold_6473/dmoj_caf060210_exons_tandy.gff | perl -n $td/genebest.perl # gene bestids scores/method predictors num score lost tandem near equal GI_BATZ_CNA_: 247 76.10 1.89 0.09 0.05 0.88 GI_BREN_NSC_: 285 74.16 1.90 0.10 0.06 0.82 GI_DGIL_SNO_: 641 78.06 1.36 0.17 0.07 0.76 GI_EISE_CEX_: 582 76.64 1.79 1.06 0.16 0.74 GI_EISE_CGW_: 637 79.19 1.55 1.29 0.12 0.84 GI_NCBI_GNO_: 339 73.55 3.44 0.24 0.14 0.76 GI_PACH_GMP_: 329 75.27 3.95 0.05 0.04 0.91 GI_RGUI_GID_: 403 70.28 3.38 0.13 0.07 0.76 GLEAN_ : 407 72.99 1.88 0.18 0.09 0.75 TRdmoj_ : 347 74.18 4.46 0.21 0.11 0.84 dmoj_GLEANR_: 352 73.30 2.20 0.22 0.11 0.74 //? GI_EISE_CEX_.-: 2 46.50 0.00 0.50 0.50 0.50 melon.% grep bestids= scaffold_6680/dmoj_caf060210_exons_tandy.gff | perl -n $td/genebest.perl # gene bestids scores/method predictors num score lost tandem near equal GI_BATZ_CNA_: 505 72.10 2.44 0.23 0.11 0.76 GI_BREN_NSC_: 642 70.87 2.42 0.29 0.14 0.72 GI_DGIL_SNO_: 1199 73.71 1.77 0.32 0.13 0.67 GI_EISE_CEX_: 1672 74.84 1.75 1.07 0.18 0.70 GI_EISE_CGW_: 2035 75.98 1.25 2.29 0.23 0.69 GI_NCBI_GNO_: 721 69.72 3.17 0.43 0.17 0.70 GI_PACH_GMP_: 832 73.49 3.67 0.38 0.12 0.81 GI_RGUI_GID_: 802 66.71 4.31 0.27 0.13 0.69 GLEAN_ : 887 70.30 1.97 0.40 0.17 0.69 TRdmoj_ : 1080 71.87 2.83 1.18 0.28 0.66 dmoj_GLEANR_: 849 69.49 2.38 0.41 0.18 0.67 #........................... melon.% grep bestids= scaffold_6541/dmoj_caf060210_exons_tandy.gff | perl -n $td/geneinfo.perl # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgns gnids maxgns methods nexons tandkno tandnew tandems GI_BATZ_CNA_ 36 5.86 3.61 6.19 33.89 70.36 16.00 6.31 7.08 13 0 0.72 GI_BREN_NSC_ 40 5.38 3.35 5.78 31.80 67.42 15.30 6.28 6.58 15 0 0.68 GI_DGIL_SNO_ 161 2.00 1.06 2.14 9.45 21.01 5.53 2.69 3.02 17 4 0.29 GI_EISE_CEX_ 2 0.00 3.50 3.50 6.00 8.00 2.00 7.00 4.00 2 0 1.00 GI_EISE_CGW_ 3 0.00 2.67 3.33 5.00 7.67 2.00 6.00 4.00 2 0 0.67 GI_NCBI_GNO_ 55 3.75 2.29 4.15 26.31 58.31 13.91 5.82 4.60 16 0 0.71 GI_PACH_GMP_ 3 0.00 2.67 3.33 5.00 7.67 2.00 6.00 4.00 2 0 0.67 GI_RGUI_GID_ 65 3.83 2.23 3.83 22.98 51.94 12.68 4.95 4.63 14 0 0.63 GLEAN_ 67 3.54 2.36 4.04 23.13 51.48 12.43 5.61 4.61 18 0 0.66 TRdmoj_ 27 3.67 2.78 3.48 16.78 28.22 8.56 2.56 4.85 9 0 0.48 all 212 1.74 0.92 1.88 7.72 17.18 4.83 2.47 2.73 24 4 0.28 dmoj_GLEANR_ 64 3.70 2.44 4.11 24.09 53.67 12.94 5.73 4.69 18 0 0.67 melon.% grep bestids= scaffold_6473/dmoj_caf060210_exons_tandy.gff | perl -n $td/geneinfo.perl # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgns gnids maxgns methods nexons tandkno tandnew tandems GI_BATZ_CNA_ 200 0.32 2.40 7.33 8.53 17.23 3.59 9.77 7.47 76 2 0.53 GI_BREN_NSC_ 226 0.31 2.26 6.85 8.02 16.39 3.40 9.49 7.00 83 2 0.52 GI_DGIL_SNO_ 455 0.20 1.22 4.49 4.31 8.84 2.24 5.33 4.66 85 4 0.30 GI_EISE_CEX_ 218 0.24 2.22 6.29 7.44 15.28 3.31 8.87 6.54 76 1 0.57 GI_EISE_CGW_ 199 0.17 2.50 6.99 8.68 17.40 3.59 10.05 7.16 79 2 0.71 GI_NCBI_GNO_ 267 0.31 2.00 5.88 7.05 14.44 3.02 8.71 6.06 86 2 0.55 GI_PACH_GMP_ 195 0.17 2.39 6.84 7.95 17.06 3.55 10.03 7.01 72 1 0.66 GI_RGUI_GID_ 316 0.29 1.70 5.41 6.11 12.95 2.90 7.67 5.59 87 2 0.47 GLEAN_ 277 0.31 1.95 6.00 7.18 14.87 3.19 8.80 6.17 89 2 0.53 TRdmoj_ 224 0.16 2.08 6.31 7.20 15.41 3.29 9.07 6.45 75 1 0.57 all 629 0.23 0.95 3.55 3.34 7.27 2.03 4.51 3.77 90 4 0.28 dmoj_GLEANR_ 257 0.32 2.09 6.21 7.68 15.70 3.33 9.18 6.37 88 2 0.58 melon.% grep bestids= scaffold_6680/dmoj_caf060210_exons_tandy.gff | perl -n $td/geneinfo.perl # gene info, all-methods scores/method predictor count eqfar eqnear eqsame fullgns gnids maxgns methods nexons tandkno tandnew tandems GI_BATZ_CNA_ 399 0.27 2.70 7.13 9.63 20.36 4.32 10.12 7.38 145 7 0.91 GI_BREN_NSC_ 466 0.28 2.56 6.58 9.00 18.98 4.23 9.53 6.84 167 7 1.00 GI_DGIL_SNO_ 784 0.23 1.59 4.82 5.64 12.18 3.03 6.33 5.05 172 11 0.64 GI_EISE_CEX_ 558 0.24 2.18 5.32 7.92 16.33 3.97 8.30 5.68 160 14 0.96 GI_EISE_CGW_ 483 0.16 2.58 6.13 9.61 19.32 4.59 9.68 6.43 164 15 1.25 GI_NCBI_GNO_ 509 0.36 2.37 6.05 8.63 17.94 4.21 9.08 6.37 163 13 0.98 GI_PACH_GMP_ 437 0.18 2.40 6.41 9.33 19.42 4.35 9.90 6.68 137 11 0.93 GI_RGUI_GID_ 612 0.27 2.12 5.46 8.07 16.68 4.07 8.38 5.73 177 12 0.99 GLEAN_ 584 0.26 2.24 5.71 8.60 17.82 4.28 8.96 5.98 178 16 1.06 TRdmoj_ 524 0.22 2.38 5.77 8.81 17.92 4.28 9.08 6.10 162 16 1.12 all 1112 0.27 1.25 3.73 4.68 10.09 2.81 5.33 4.03 188 19 0.61 dmoj_GLEANR_ 572 0.27 2.29 5.69 8.68 18.06 4.33 9.05 5.99 178 16 1.09 // GI_EISE_CEX_.- 2 1.50 4.00 7.00 11.00 30.50 6.00 10.50 7.00 2 0 2.50 =item togff2 test ggb # #!/usr/bin/perl # $v=9; # while(){ # s/tandy\d/tandy$v/g; s/tandy\.\w+/tandy$v/g; # s/^/#/ if(m/match\t/); # s/(ID|Parent)=([^;]+);/gene "$2" ; /; # s/tclass=(\w+);/Note "$1" ; /; # s/;cid=.*$//; # s/\t[\d\.]+e-/\t/; # print; # } # # __DATA__ ##gff-version 2 #reference = scaffold_4:939037-957688 [tandy6] feature = alignment:tandy6 region:tandy6 #bgcolor = #A0A0C0 glyph = graded_segments height = 6 connector = dashed box_subparts = 0 =cut =item gene groups and altexons These 3 gene groups share exons via near-same (altexons); should they become 1 group? run a risk of making huge groups via shared exons # query: scaffold_4/dpulex1_exons.nr; genome: scaffold_4/dpulex1.fa # Location: scaffold_4:830000-1030000 scaffold_4 tandy.common match 939510 952854 . - . ID=td_c. Dappu1_FM5_191620;tdclass=common;cid=6;ids=DP_DGIL_SNO_00002554[42, 12,9],Dappu1_FM5_191620[2,2,0],NCBI_GNO_454044[3,2,1], NCBI_GNO_456044[9,4,2],NCBI_GNO_458044[13,3,2];nexons=85;altexons=48 scaffold_4 tandy.common match 944216 957688 . - . ID=td_c. DP_DGIL_SNO_00002556;tdclass=common;cid=7;ids=Dappu1_FM5_42415[1,1,0 ],NCBI_GNO_460044[3,3,0],NCBI_GNO_462044[2,2,0];nexons=30;altexons=1 scaffold_4 tandy.common match 939037 950039 . - . ID=td_c. NCBI_GNO_452044;tdclass=common;cid=8;ids=DP_DGIL_SNO_00002553[2,1,1] ,NCBI_GNO_452044[7,5,1];nexons=20 =cut