We implemented a parquet writer in DataHaskell/Dataframe. Using it is simple; you need only pass your dataframe into the writeParquet function which writes a...
Starting with v2.0, scheduled for fall 2026, DuckDB will support asynchronous reads of Parquet and CSV files. This can significantly speed up queries when synchronous I/O does not saturate the available bandwidth, as is typical in EC2/S3 compute-storage…
I built jq for Parquet, then ghosted three v0.14 issues for 21 days. Here's how a hashtag, Copilot, and Cursor finally shipped them — and what I learned when DuckDB lied to me on the cover image.
A Quick Guide to Reading and Writing Files in PySpark How to Read and Write CSV, Parquet, and JSON Files in PySpark. From the previous blogs about PySpark, you had some idea of why you need to …