TL;DR
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
Polars 2.0 makes its streaming engine the default for LazyFrame collection and enables initial out-of-core processing that can spill supported operations to disk. The project also expands SQL support and reports favorable TPC-H and TPC-DS benchmark results, while noting limitations in its tests and that some operations, including joins and group-bys, cannot yet spill.
Polars has shipped version 2.0, making its streaming engine the default when users collect a LazyFrame and enabling initial spill-to-disk support for some operations. The release also expands SQL support and adds a Map data type, changes that affect how data workloads are executed and how users may need to handle row order and memory limits.
Under the new default, calling collect on a LazyFrame uses the streaming engine. Polars says this can improve memory use and performance on many queries. However, the engine does not preserve observable row order by default for certain operations, including joins, group-bys and unpivoting. Users who need ordering for those operations can set maintain_order=True.
Version 2.0 also enables out-of-core processing, which lets supported operations spill data to disk when memory use reaches roughly 80% of RAM. The default disk budget is 64 GB. The release announcement lists sort, window functions and many expressions among operations that can use the feature; joins and group-bys are not supported for spill-to-disk yet, though Polars says it plans to add them.
The release introduces a native Map dtype for Arrow MapType data, representing key-value collections in a dictionary-like form. Previously, Polars read that Arrow type as a list of structs. The project also cites stricter dtype handling and greater explicitness as changes intended to provide faster feedback during development.
Streaming Changes Query Execution
The shift to streaming by default may change both the resource profile and behavior of existing Polars workloads. The project says the change can bring memory and performance improvements on many queries, while the row-order caveat means some users may need to review queries whose results depend on ordering. The explicit maintain_order=True option gives users a way to request that behavior where supported.
Disk spilling can help queries that exceed available memory finish rather than relying solely on RAM, but the current support is partial. Since joins and group-bys cannot yet spill, the feature is not a general guarantee that any oversized workload will complete. The 64 GB disk budget and approximate 80% RAM threshold are defaults stated by the project, which says the threshold may need tuning.
For teams evaluating Polars as a SQL engine, the release’s benchmark results offer a point of comparison, not an independent verdict. The tests were designed and published by Polars, and results depend on hardware, workload and benchmark methodology. Readers should treat the reported performance as the project’s measurements and reproduce them against their own queries before making deployment decisions.
high performance data processing laptop
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
SQL Benchmarks and Release Scope
Polars describes 2.0 as a major version change tied in part to the new streaming default and its row-order implications, rather than solely as a feature-count milestone. The project says it has expanded SQL coverage and improved its optimizer and engine, citing join reordering, common-subplan elimination and dynamic predicates or bloom filters.
To compare SQL performance, Polars ran queries based on TPC-H and TPC-DS against DuckDB 1.5.6, DuckDB 2.0 alpha and DataFusion 54.0.0. Tests used two AWS machine types, repeated queries five times in a hot setting and compared the best run. Polars reports that it was fastest in all but one benchmark under its default configuration, while limiting Polars to 32 cores was competitive or winning across the reported benchmarks.
The benchmark had exceptions: DataFusion timed out on TPC-DS query 72, timed out once on query 67 and ran out of memory on TPC-H query 18 on the smaller machine. Those queries were excluded from results for every engine. Polars also reported that its default configuration had overhead on the 192-thread machine that hurt smaller queries, and said it had diagnosed the issue and hoped to fix it in a later release. It published a repository for others to replicate the tests.
“Calling collect on a LazyFrame will now default to the streaming engine.”
— Polars, in its Polars 2.0 release announcement
As an affiliate, we earn on qualifying purchases.
Limits of Current Spill Support
The announcement does not give a release calendar date, so the timing is identified here only as a shipped release. It also does not quantify a general performance gain across users’ workloads. Polars’ benchmark claims are based on its disclosed setup; independent replications and results on other hardware or datasets are not included in the supplied material.
It remains unclear when out-of-core support for joins and group-bys will arrive, how much the RAM threshold may need adjustment in different environments, and how the new defaults will affect individual applications. Users should verify ordering requirements and test memory behavior on their own workloads rather than assume every operation can spill to disk.
As an affiliate, we earn on qualifying purchases.
More Out-of-Core Operations Planned
Polars says it plans to extend spill-to-disk support to joins and group-bys, but the release announcement gives no date for that work. The project also says it hopes to address the performance overhead it observed when running on 192 threads in a later release.
For now, users moving to 2.0 can check whether queries rely on row order, use maintain_order=True where needed, and test the enabled streaming and spill behavior with representative data. Those steps can help establish whether the new defaults suit a particular workload while the project develops the announced improvements.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main change in Polars 2.0?
LazyFrame collection now defaults to the streaming engine. The release also enables initial spill-to-disk support and adds SQL and datatype changes.
Does the streaming engine preserve row order?
Not by default for some operations, including joins, group-bys and unpivoting, according to Polars. Users who need observable ordering can set maintain_order=True.
Which operations can spill data to disk?
The announcement lists sorts, window functions and many expressions as supported. Joins and group-bys are not supported yet; Polars says those are planned additions.
Are the performance benchmark claims independently verified?
The figures described in the announcement come from Polars’ own tests of TPC-H and TPC-DS workloads against DuckDB and DataFusion. Polars shared a benchmark repository for replication, but the supplied source does not include independent results.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
