upvote
Polars relies on threading heavily, even when streaming files. And it appears that each thread loads quite a bit of memory. I've encountered OOM issues when incrementally reading Arrow IPC files which had very large batches. Fixed it by setting $POLARS_MAX_THREADS to 1, which amusingly also improved the performance on my very narrow task.
reply
Peak memory usage per query is in the raw data (in `results/`) in the repository: https://github.com/pola-rs/polars-2.0-benchmark/.
reply