We have a retrieval benchmark based on RULER which we've been using to ensure that the model maintains complete awareness of the full context window.
All our benchmarks are open source so you can check it out here if you'd like:
https://github.com/magnitudedev/magnitude/blob/main/inferenc...