Dear Dr. Ou and EDTA Development Team,
I hope this message finds you well.
I am currently using EDTA for transposable element annotation and downstream evolutionary landscape analysis in my genome project, and I sincerely appreciate the excellent work your team has done in developing this pipeline.
I have a question regarding the calculation of divergence values for DNA transposons, especially TIR elements, in the EDTA.TEanno.bed output.
From the annotation results, I noticed that for structurally annotated TIR elements (method=structural), the “identity” field appears to be derived from structural features such as TIR/TSD similarity, rather than from sequence alignment to consensus sequences. In downstream scripts (e.g., div_table.pl), this identity is converted to divergence using:
div = 100 × (1 − identity)
My concern is whether this structural identity can be interpreted as a reliable proxy for sequence divergence from consensus, and thus used for insertion age estimation in TE landscape analyses, in the same way as Kimura distances from RepeatMasker.
Specifically, I would like to ask:
Is the identity value for structurally annotated TIR elements intended to approximate evolutionary divergence from consensus?
If not, would you recommend using only RepeatMasker-based homology annotations for TE age estimation?
Are there any best-practice recommendations for handling DNA transposons in EDTA when generating TE divergence landscapes?
I would be very grateful for your guidance on this issue, as it is important for the interpretation of TE dynamics in my study.
Thank you very much for your time and consideration.
Dear Dr. Ou and EDTA Development Team,
I hope this message finds you well.
I am currently using EDTA for transposable element annotation and downstream evolutionary landscape analysis in my genome project, and I sincerely appreciate the excellent work your team has done in developing this pipeline.
I have a question regarding the calculation of divergence values for DNA transposons, especially TIR elements, in the EDTA.TEanno.bed output.
From the annotation results, I noticed that for structurally annotated TIR elements (method=structural), the “identity” field appears to be derived from structural features such as TIR/TSD similarity, rather than from sequence alignment to consensus sequences. In downstream scripts (e.g., div_table.pl), this identity is converted to divergence using:
div = 100 × (1 − identity)
My concern is whether this structural identity can be interpreted as a reliable proxy for sequence divergence from consensus, and thus used for insertion age estimation in TE landscape analyses, in the same way as Kimura distances from RepeatMasker.
Specifically, I would like to ask:
Is the identity value for structurally annotated TIR elements intended to approximate evolutionary divergence from consensus?
If not, would you recommend using only RepeatMasker-based homology annotations for TE age estimation?
Are there any best-practice recommendations for handling DNA transposons in EDTA when generating TE divergence landscapes?
I would be very grateful for your guidance on this issue, as it is important for the interpretation of TE dynamics in my study.
Thank you very much for your time and consideration.