I have a compressed CSV file compressed as csv.gz which I want to run some processing on. I generally go with Polars because it is more memory-efficient and faster. Here is the code which I am using to lazily read and filter it before I can run some other processing on it.
df = (
pl.scan_csv(underlying_file_path, try_parse_dates=True, low_memory=True)
.select(pl.col("bin", "price", "type", "date", "fut"))
.filter(pl.col("date") == pl.col("date").min())
.collect()
)
On running this, I seem to run out of memory because I just get a Killed message with no other output. On the other hand, when I try to read and print the same dataframe with Pandas:
df = pd.read_csv(underlying_file_path, usecols=["bin_endtime", "strike_price", "opt_type", "expiry_date", "cp_fut"], parse_dates=True, low_memory=True)
This works fine and I am able to print and process the file fine. This is uncanny because up till now, I've always noticed that Polars is able to handle larger data than Pandas and is faster while doing so. Why could this be happening?
Details
- OS: Ubuntu 22.04.5 LTS
- Pandas Version: 2.3.3
- Polars version: 1.35.2
- Python version: 3.10.12
- File size: 2.1G
- Number of rows in the CSV file: 42.39M
Please let me know if any other details are required.