undefined

points

[-]

Would be still interesting whether it degraded the performance in that case. Further, many non-agentic benchmarks consist of many short tasks, so one could fill the context with task/response pairs from other tasks (like in a standard chat environment) and then ask the current task at the end. Given that the tasks are probably somewhat similar, context rot should occur.