
Polars 2.0 shipped today, and the headline change is deceptively quiet. Call .collect() on a LazyFrame and you are now running on the streaming engine. No config flag, no migration step. Ritchie Vink, Polars creator, framed it simply: “your code will be faster without doing anything.” The catch — and there always is one — is that row order is no longer guaranteed. If your pipeline has been silently relying on join or group-by order, you do not have a crash waiting to happen. You have wrong results waiting to happen.
The Change That Can Corrupt Your Data Without Warning
In Polars 2.0, join, group_by, and unpivot operations no longer preserve row order by default. The streaming engine processes data in parallel chunks, and those chunks do not reassemble in the order you might expect. Your numbers stay arithmetically correct. Your code does not throw an exception. But entities can attach to the wrong rows, and if you are feeding results into a downstream system that assumes consistent ordering, you are now in production data corruption territory.
The fix is one flag: maintain_order=True on affected operations. The harder part is knowing which operations in a six-month-old pipeline need it. Polars recommends a codebase audit before upgrading anything in production. The official migration guide covers the full scope.
# Before 2.0: join silently preserved row order
result = df.lazy().join(other, on="id").collect()
# After 2.0: row order not guaranteed — add this if order matters
result = df.lazy().join(other, on="id", maintain_order=True).collect()
# Opt back to in-memory engine entirely while you audit
pl.Config.set_engine_affinity("in-memory")
What You Actually Get From the Upgrade
The streaming default is not just a performance optimization. It is an architectural shift. The old engine loaded entire datasets into RAM before processing. The new one works in CPU-cache-optimized chunks, which changes the economics of large-scale data work significantly.
Out-of-core processing comes with it. When RAM hits roughly 80 percent utilization, Polars spills to disk automatically, with a 64 GB default disk budget. Sorts, window functions, and most expressions are covered. Joins and group-bys remain in-memory for now — that is on the 2.x roadmap.
Polars 2.0 also ships a native Map dtype, replacing the List(Struct) workaround for dictionary-like data. You get .get(), .keys(), .values(), and contains_key() as first-class operations. SQL support got a meaningful upgrade too, with join reordering, bloom filters, and common-subplan elimination added to the query optimizer.
The Benchmark Numbers (Honest Edition)
Polars’ own TPC-H and TPC-DS benchmarks show it outperforming both DuckDB and DataFusion on most queries at scale. DataFusion timed out on two TPC-DS queries and ran out of memory on TPC-H. That is not a great look for DataFusion.
For day-to-day Python work, the numbers are sharper. On H2O.ai’s group-by benchmark at 10 million rows, Polars finishes in 0.45 seconds. Pandas takes 12.5 seconds — a 27x gap, not a rounding error.
The honest DuckDB comparison is more nuanced. DuckDB still wins on memory efficiency by a significant margin — 1.3 GB versus Polars’ 17 GB on a 140 GB file. Polars wins CSV reads and joins; DuckDB leads on window functions; group-bys are roughly even. If you are memory-constrained, DuckDB holds its ground. A mixed Polars-plus-DuckDB stack, leveraging Arrow’s zero-copy buffer sharing, can outperform either library alone on wide table scans.
Should You Switch From Pandas?
The honest answer depends on what you are building. Pandas is still at 42 percent adoption among professional Python developers. Polars is at 11 percent. It is not replacing anything in the short term.
What is shifting is where new work goes. Enterprise data teams at Deutsche Bahn, Citizens Bank, and Rabobank have published migration case studies with reported speedups of 10x to 20x and cloud bill reductions around 25 percent. Polars job postings on LinkedIn are up 450 percent year-over-year. The migration is real, just not complete.
The practical rule: start new pipelines in Polars. Migrate legacy code selectively, after auditing for row-order dependencies. Do not mass-upgrade a working Pandas codebase on a Friday afternoon to chase benchmark numbers.
API Changes to Fix Before You Upgrade
Beyond row order, several APIs renamed or disappeared in 2.0:
melt()→unpivot()read_csv()→scan_csv().collect()LazyFrame.profile()removedjoin_nullsparameter →nulls_equal
New exception types — AttributeRemovedError and ArgumentRemovedError — will surface anything you missed on first run. The full deprecation list and upgrade walkthrough are in the official docs. Run your test suite before touching production.













