upvote
If Pandas was the great grandfather, R data.frame is the great-great grandfather. R data.frames directly inspired Pandas.
reply
Agreed. For its time, R was a lot of fun to play with data pipelines, visuals, Quarto, and frontier stats. It's a shame it's so hard to make it work for a large swath of production use cases.
reply
Do you have a take on when each of these three choices is the best one? I totally agree that these are the good choices, but I still find myself hesitating about which thing to reach for when!
reply
Depends on need. We started using PyArrow on a reporting microservice when we realized we needed no additional functionality that pandas provided since it has better data type ergonomics. DuckDb is a great go-to for SQL based transformations when working with parquet files outside a managed system like Databricks. I want to actually test duckdb versus polars with a few lower level places like iceberg on S3.
reply
Yeah makes sense. The SQL thing has also been my differentiator (and yeah, straight up arrow if you aren't doing much transformation of the data), but now I'm curious whether polars sql might be just as good. I kind of like that duckdb allows me to work with a database file, like a sqlite db. But maybe persisting parquet (or arrow directly?) is just as good?

This is why I asked someone else who is also figuring this out!

reply
Really comes down to query speed and management cognitive cost. I look forward to trying the new version for polars.
reply