each individual benchmark that is combined into agentic coding index was compared on xhigh between ua / codex / pi, and headline improvement was calculated on xhigh, but then agentic coding index pareto chart by default includex codex max, hense the confusion.
[features.multi_agent_v2]
enabled = true
wait_agent_enabled = true
min_wait_timeout_ms = 10000 # 10s
default_wait_timeout_ms = 300000 # 5m
max_wait_timeout_ms = 3600000 # 1h
Seems to do the job and reduce usage; I just ran Astra for ~5 hours (using a goal) and it used the last 30% of my usage. And now they released GPT-6 Sol and Luna (which is basically 5.6 Sol and Luna, but a bit better and also 50% cheaper) ;_;