upvote
From recent Python Bytes podcast (https://pythonbytes.fm/episodes/show/496/a-lake-house-in-sea...)

> 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory

Python Vs Rust : In terms for speed - No comparison

(The above episode transcript has a link to blog post titled "Pandas should go extinct" )

reply
I've nearly entirely switched to DuckDB for anything more than like 500 or 1,000 rows or if there are a tonne of columns.

Polars is great, but I'm just too used to the Pandas API to use it as a replacement for the cases where DuckDB is overkill.

reply
That's a shame, because Pandas has a really quirky/legacy-burdened API and Polars is super clean by comparison. As someone who had Spark and Pandas experience before switching, Polars felt like Pyspark without the added mental overhead of needing you to think about multi-worker-node parallelism
reply
It's just muscle memor, I've been using Pandas daily for over a decade, but I'll likely eventually switch over to Polars.

Part of the problem, as I said, is that I'm just spoiled by DuckDB when performance matters.

reply
Polars is effectively a full replacement for Pandas for 99.9% of all cases. The only exception I'm really aware of is if you're working with geospatial data, as there isn't yet a "Geopolars" equivalent of the commonly used "Geopandas". However, Geopolars is still in active development and should eventually be production ready.
reply
There is a geospatial package for duckdb though which is pretty slick. It is actually how I first learned of duckdb. We were dealing with nationwide parcel datasets and need to apply transformations nationwide and save out to more parquet files. It was easier and cheaper to replace all of the pandas workloads with duckdb.
reply
it’s on the correct path. i use rust for geo spatial and the gap with c, c++ closing rapidly or negligible in most cases

from https://github.com/pola-rs/geopolars/tree/main

Comparison with GeoPandas

Imitation is the sincerest form of flattery! GeoPandas — and its underlying libraries of shapely and GEOS — is an incredible production-ready tool.

GeoPolars is nowhere near the functionality or stability of GeoPandas, but competition is good and, due to its pure-Rust core, GeoPolars will be much easier to use in WebAssembly.

reply
Note that (for the time being), we started geopolars developement under: https://github.com/pola-rs/geopolars/tree/dev
reply
> Note that (for the time being), we started geopolars developement under: https://github.com/pola-rs/geopolars/tree/dev

This is awesome!! I'd looked at the project only a month or so and it appeared abandoned, but I must have missed the off-main-branch development going on!

reply
Pandas is better for slight in-place or per-row modifications, for loading from less conventional datasets, for transposition/more nuanced row-based aggregation/multi-axis manipulation, for performance when multiprocessing can be used, for interoperability with other libraries (e.g. plotting and statistics)... When you need to operate on huge datasets, use DuckDB, because its performance is even now on-par with Polars regarding speed, while handling huge or more complex joins is a huge win for DuckDB because it better offloads intermediate results to disk, while Polars just dies on me. I haven't tried such joins in Polars 2 though
reply
We were able to do a full port of GFQL from pandas to polars, cypher graph queries on dataframes, including both our CPU + GPU modes, and hit massive speedups: https://www.graphistry.com/blog/cypher-on-polars-cpu-gpu-gra...

It's been impressive!

reply
Yes and no, its not replacing the reason why pandas was popular ie data scientists, but it a full replacement of its pipeline usage, And I would saw also beating out spark
reply
my understanding is Polars is faster, scales better without using external solutions, better API, +Rust. Pandas wins if you want to use what the vast majority of folks are using and have used in the past. Probably has a more complete set of helpers / recipes for the little things you bump into when using it thoroughly, but in the age of LLMs, I think that's minor.
reply
>Pandas wins if you want to use what the vast majority of folks are using

Vast majority of skilled developers are now using Polars, unless they are constrained by lack of Narwhals support in their third-party library of choice (e.g. Great Expectations, SHAP). That's the more important trend to follow.

reply
It’s the API that gave me the push to leave Pandas. 10 or more years of occasional Pandas use and I still had to google for any non-trivial queries.

In that regard, I’m still waiting for a credible jq replacement…

reply
Learn SQL and interface with Duck. You will be 100x faster than Pandas/Polaris duo at fraction of memory. Also SQL is supported literally everywhere with a much more capable than Pandas API. Duck outputs to a Pandas Dataframe, but just treat that like a dictionary. Do all your processing, filtering and aggregation in Duck.

Also, try fx.wtf as a replacement for jq. it comes with a in-built tui viewer that supports vi-keybindings. Ecmascript is built into fx.wtf so you can query the JSON with JS notation (where JSON was born). You can use any JS functions including map/reduce/filter or perform any kind of transformation instead of learning jq dsl that you will forget tomorrow.

reply
Polars also supports SQL now, per the announcement we're all commenting on.
reply
Avoiding pandas developers is a great reason to use polars imo
reply
It has been for me. I greatly prefer the API, it fits my mental model much better. Give it a try!
reply
See "Pandas should go extinct": https://news.ycombinator.com/item?id=49668198

tl;dr yes

reply
It kind of reminds me of Dungeons and Dragons, in a sense. Pandas is so popular that people learn/use it because of its popularity, rather than because it's the best for any one use-case. Polars is cleaner, faster, and can easily be scaled up for production use-cases so your little POC script for local analysis can be productionized really easily, but there are holdouts still on Pandas because it was so complicated to learn with so many little extra rules to learn to avoid paper cuts that they feel like learning another data manipulation tool would be really hard.

DnD does the same thing: it's the most popular but its rules are this awkward hybrid of legacy cruft and some modern ideas, so learning it a huge effort, which means most people who play it aren't willing to try any other RPG systems even though most of them are dramatically easier to learn because they were built with a clean design from the ground-up.

It's the sunk-cost fallacy as applied to learning something complex, combined with something like the horn effect (inverse of the halo effect) making any competitors look equally complex even if they're not, causing long-time Pandas users/DnD players to strongly resist even looking at other options. Basically the frustration of learning these older systems seemingly traumatizes some people into never straying.

reply
There are awkward things. For example, if you ingest a nanosecond resolution timestamp, there's no way to re-export that out of the Polars dataframe with nanosecond resolution.
reply
I'm not sure this is true e.g. you can specify schema `pl.Datetime("ns")`. That survives roundtrip to/from parquet in my experience. True though that the default is `us` whereas pandas defaults to `ns`.

https://docs.pola.rs/api/python/stable/reference/expressions...

reply
Feels like a feature to me. You could just keep the original timestamp as a string.
reply
From experience I'll say one thing that Polars doesn't have great interfacing with is doing things like string concatenation, like taking multiple columns and combining them in with static string text in complex ways to create new columns.

Pandas has a really simple ability to just define a new column with

`df['col_a'] + "text" + df[col_b']` where "text" can be any string text inbetween your column values from col_a and col_b

If i remember correctly while you can do pl.col("col_a") + pl.("col_b") for plain concatenation, you can't mix in static text strings like you can with pandas and I haven't found really elegant ways to do that personally. Whereas I've found polars doesn't have as simple of a way to do that. You can choose to add one separator and make that anything you want, but only one separator and only inbetween the two values (so no suffixes or prefixes for example).

That being said, I hate everything to do with the pandas API (especially with its indexing system) and really prefer the more polars API for anyone coming from a SQL or database background. Pandas really shows its sort of academia background rather than a data engineering origin.

reply
I'm not sure what you tried. Is this not what you want?

    >>> df = pl.DataFrame({"x": ["a", "b", "c"], "y": ["d", "e", "f"]})
    >>> df.with_columns(new=pl.col.x + " text " + pl.col.y)
    shape: (3, 3)
    ┌─────┬─────┬──────────┐
    │ x   ┆ y   ┆ new      │
    │ --- ┆ --- ┆ ---      │
    │ str ┆ str ┆ str      │
    ╞═════╪═════╪══════════╡
    │ a   ┆ d   ┆ a text d │
    │ b   ┆ e   ┆ b text e │
    │ c   ┆ f   ┆ c text f │
    └─────┴─────┴──────────┘
reply