TIS 2.0: Token Importance Scoring Now Eliminates Position Bias in RAG

I currently want to use the Jina embedding model, which can embed entire documents at once. And then I noticed that there was too much noise and that tokens that are close to each other had a strong influence and probably also the tokens at the beginning and end of the document were given more weight. I tried using “distance-residual attention” and your technique could perhaps reduce the noise even more:

https://huggingface.co/proxy/discuss.huggingface.co/t/attention-weighted-memory-graphs/178271/3