It's like you get a really good 'query planner' like a DB would give you, but for your notebooks/scripts/etc. Much better than pandas imo.
DuckDB is an in process OLAP, I’ve been using it a lot, and I am keen to use Polars but DuckDB seemings to be flying for me.
> 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory
Python Vs Rust : In terms for speed - No comparison
(The above episode transcript has a link to blog post titled "Pandas should go extinct" )
Polars is great, but I'm just too used to the Pandas API to use it as a replacement for the cases where DuckDB is overkill.
Part of the problem, as I said, is that I'm just spoiled by DuckDB when performance matters.
from https://github.com/pola-rs/geopolars/tree/main
Comparison with GeoPandas
Imitation is the sincerest form of flattery! GeoPandas — and its underlying libraries of shapely and GEOS — is an incredible production-ready tool.
GeoPolars is nowhere near the functionality or stability of GeoPandas, but competition is good and, due to its pure-Rust core, GeoPolars will be much easier to use in WebAssembly.
This is awesome!! I'd looked at the project only a month or so and it appeared abandoned, but I must have missed the off-main-branch development going on!
It's been impressive!
Vast majority of skilled developers are now using Polars, unless they are constrained by lack of Narwhals support in their third-party library of choice (e.g. Great Expectations, SHAP). That's the more important trend to follow.
In that regard, I’m still waiting for a credible jq replacement…
Also, try fx.wtf as a replacement for jq. it comes with a in-built tui viewer that supports vi-keybindings. Ecmascript is built into fx.wtf so you can query the JSON with JS notation (where JSON was born). You can use any JS functions including map/reduce/filter or perform any kind of transformation instead of learning jq dsl that you will forget tomorrow.
tl;dr yes
DnD does the same thing: it's the most popular but its rules are this awkward hybrid of legacy cruft and some modern ideas, so learning it a huge effort, which means most people who play it aren't willing to try any other RPG systems even though most of them are dramatically easier to learn because they were built with a clean design from the ground-up.
It's the sunk-cost fallacy as applied to learning something complex, combined with something like the horn effect (inverse of the halo effect) making any competitors look equally complex even if they're not, causing long-time Pandas users/DnD players to strongly resist even looking at other options. Basically the frustration of learning these older systems seemingly traumatizes some people into never straying.
https://docs.pola.rs/api/python/stable/reference/expressions...
Pandas has a really simple ability to just define a new column with
`df['col_a'] + "text" + df[col_b']` where "text" can be any string text inbetween your column values from col_a and col_b
If i remember correctly while you can do pl.col("col_a") + pl.("col_b") for plain concatenation, you can't mix in static text strings like you can with pandas and I haven't found really elegant ways to do that personally. Whereas I've found polars doesn't have as simple of a way to do that. You can choose to add one separator and make that anything you want, but only one separator and only inbetween the two values (so no suffixes or prefixes for example).
That being said, I hate everything to do with the pandas API (especially with its indexing system) and really prefer the more polars API for anyone coming from a SQL or database background. Pandas really shows its sort of academia background rather than a data engineering origin.
>>> df = pl.DataFrame({"x": ["a", "b", "c"], "y": ["d", "e", "f"]})
>>> df.with_columns(new=pl.col.x + " text " + pl.col.y)
shape: (3, 3)
┌─────┬─────┬──────────┐
│ x ┆ y ┆ new │
│ --- ┆ --- ┆ --- │
│ str ┆ str ┆ str │
╞═════╪═════╪══════════╡
│ a ┆ d ┆ a text d │
│ b ┆ e ┆ b text e │
│ c ┆ f ┆ c text f │
└─────┴─────┴──────────┘Everything that I build greenfield moving forward I plan to use DuckDB, Polars, or PyArrow. Pandas was a great grandfather of a project (I actually cut my OSS contrib teeth on it, how the time flies)! I'll always appreciate the improvement pandas brought over SAS.
This is why I asked someone else who is also figuring this out!
When you see a blog post like this, never interpret it as “database A is X% faster than database B”, there are just too many factors.
It’s more like “we put focused work into performance improvements and we expect certain workloads to perform better than the previous release”.
This seems like a good project, and benchmarking is a good way for a development team to iterate on performance. Just want to get my take out there.
I've spent years of my life running TPC benchmarks. They're useful guidelines, but they're simply not diverse enough to show the whole picture unless your picture is very simple. Certainly no single TPC benchmark in isolation. Maybe you can get a better idea by running all of them.
The largest gaps are around non-relational style data, JSON querying and the like. Nearly every organisation has something in that space now, and TPC-H/TPC-DS don't touch it at all.
sure, nothing is 100%. But 90% may be good enough.
> JSON querying and the like. Nearly every organisation has something in that space now, and TPC-H/TPC-DS don't touch it at all.
unnesting json into relational data is some trivial op, and then you come back into tpc-h/ds realm.
I mean, everything is representative if you just phrase the bits you don't want to do as "some trivial op". By this logic, Clickbench is a great and totally representative benchmark because joins are also trivial ops.
Billions of dollars are spent processing semi-structured data. It's a critical part of most analytical pipelines. You can't just ignore it and pretend like your benchmark is representative.
I tried my best to be as transparent and fair as possible, running everyone with out-of-the-box settings on a third party's queries (DuckDB), a third party's data generator (tpcgen-rs, from the DataFusion guys), on a stock setup available to everyone (AWS machines).
The one exception is that we also ran Polars locked to 32 threads on the large machine (in addition to the out-of-the-box setup), which was to highlight we can do a lot better on small data. We still suffer from a relatively high constant overhead on very high CPU count machines if the data isn't large enough, but I'm working on fixing that. It's possible that DuckDB / DataFusion have similar scaling issues with high thread counts and would do better with 32 threads as well, I didn't test that.
I think it is over-complication. TPC results used to be reported on some standard machines you can order, and now every one uses AWS metal for that. Its Ok if project has configs tuned to popular machine, I think it is representative approach.
Happy that I can upgrade to 2.0 final tonight.
For example our join currently does a full partition into T partitions, for each of the T threads. Overall we create T^2 partitions, which on a 192-core machine is non-trivial. Great if you have a ton of data to feed that with, but if you 'only' have a few dozen million rows it becomes rather small. This is the primary reason we saw in the benchmarks that Polars pinned to 32 threads beats 192 thread Polars at SF=10.
I'll be working on improving that soon. I expect that to have a big impact on SF=10, and a decent impact on ClickBench, which sits between SF=10 and SF=100 in terms of rows.
Now it seems they want to go head to head with DuckDB.
I’m guessing most of the this disparity should be solved by the steaming engine.
From a quick check our first PRs were merged to the 2.0 branch in June:
2026-06-17T21:27:51Z #27993 chore: Stop coercing `pl.col(...)` to selector ...
2026-06-18T14:19:26Z #27996 chore!: Replace multi-seed hash API with a single seed
2026-06-19T07:05:11Z #27991 chore(python!): Remove `Expr.flatten` functionApparently about something ))
I mean, it's very clear that the projects have a lot in common, even if building one on top of another directly is not a good idea.
I don't want polars to depend on on the datafusion crate, that is a complexity that is not needed.Narwhal is still fairly new, but I expect its usage to spread since most packages only require rudimentary dataframe manipulation (set a value, math been these two columns, etc) where the limited api surface is not a problem.
Narwhals is also a much cheaper dependency to add than polars/pandas/etc so it is a somewhat easy sell to incorporate.
It's also got much better support for more complex array shapes (e.g. each row storing an array). At least it did last time I used pandas!
cant wait to upgrade my Quant trading bot
when I started, I tried mimicking a human trader. if you want to expand into full quant, be aware of overengineering (past mistake of mine).
you can copy or mimic institutional desk strategies. most of the concepts are fine, but the devil is in the details.
sometimes you open too early or too late. finding that sweet spot, where you don’t want a lot of drawdown, requires a lot of fine-tuning.
more predictable
Here is an example [1] of visualizing college football games. Here are all the queries, and semantic model that power all the visualizations [2] Here is the AI generated typescript/react that does the visualizations [3]. The Malloy ecosystem has Malloyyo and Publisher which are replacements for PowerBI and Tableau and Looker. Here is another example for visualizing global trade [4].
[1] - https://mrtimo.github.io/cfb-games/games-2026.html?week=Week... [2] - https://github.com/mrtimo/cfb-games/blob/main/drives.malloy [3] - https://github.com/mrtimo/cfb-games/blob/main/dashboards/gam... [4] - https://tradeexplorer.org/