Hi Evo2 team,
When interpreting Evo2 embeddings with SAE, we consistently observe a strong activation spike within the first ~15 bp of input sequences. This pattern appears across multiple datasets, species, and SAE features.
Is this expected behavior? Could it be related to positional embeddings, sequence boundary effects, special tokens, or the SAE training procedure?
Any insights would be greatly appreciated.
Thanks!

Hi Evo2 team,
When interpreting Evo2 embeddings with SAE, we consistently observe a strong activation spike within the first ~15 bp of input sequences. This pattern appears across multiple datasets, species, and SAE features.
Is this expected behavior? Could it be related to positional embeddings, sequence boundary effects, special tokens, or the SAE training procedure?
Any insights would be greatly appreciated.
Thanks!