Guides
Practical, tool-agnostic guides on working with Parquet files. RSS
How to Append Rows to an Existing Parquet File
Parquet files are written once, so appending means adding a row group or a new file. fastparquet append=True, a DuckDB rewrite and the dataset-directory pattern.
How to Check If a Parquet File Is Corrupted (No Install Needed)
Validate a Parquet file in your browser: check the footer and schema instantly, then scan every row group with one SQL query. No pyarrow, no Spark.
Combine Multiple CSV Files into One Parquet File
Turn a folder of CSV exports into a single compressed Parquet file with one DuckDB command, handling schema drift — or do it entirely in your browser.
How to Compare Two Parquet Files and See What Changed
Diff two Parquet files by key to find added, removed and changed rows — in your browser, with DuckDB SQL, or with a short pandas script.
How to Convert an Excel File (XLSX) to Parquet
Get a spreadsheet into Parquet without pandas: export the sheet as CSV, convert it in your browser, and fix the types Excel mangled on the way out.
Batch Convert a Folder of Parquet Files to CSV
Convert every Parquet file in a directory to its own CSV in one go — with a DuckDB shell loop, PowerShell, a pyarrow script, or a multi-file drop in the browser.
Convert Parquet to CSV from the Command Line
Three ways to turn a Parquet file into CSV from a terminal — a DuckDB one-liner, Python with pandas or pyarrow — plus a no-install browser fallback.
Convert Parquet to Excel (XLSX) with DuckDB or Python
Produce a real .xlsx workbook from a Parquet file — dates stay dates, types survive — with a DuckDB one-liner or pandas, plus fixes for the two errors you will hit.
DuckDB Parquet Viewer Options Compared: CLI, UI, and Browser
DuckDB can view Parquet files through its CLI, its local web UI, Python, or a browser-based tool. Here is how each works and when to pick one.
Fix fastparquet Install and Import Errors
pip install fastparquet fails to build a wheel, or importing it raises a numpy binary incompatibility. What each error means and how to get to a working reader.
fastparquet ParquetFile: Read Columns, Filters and Row Groups
Use fastparquet's ParquetFile to inspect a file, load only the columns you need, prune row groups with filters and iterate in chunks to keep memory flat.
fastparquet vs pyarrow: Which Python Parquet Engine to Use
fastparquet and pyarrow both read and write Parquet from pandas, but they differ on speed, type support and maintenance. Here is how to choose.
fastparquet write(): Compression, Partitioning and Row Groups
Write Parquet with fastparquet that other tools read correctly: per-column compression, partition_on with hive layout, row group size, object_encoding and INT96.
How to Find and Remove Duplicate Rows in a Parquet File
Count exact duplicates, find repeated business keys and keep only the newest row per key — with DuckDB SQL in your browser or a short pandas script.
How to Flatten Nested Parquet Columns to CSV or JSON
Parquet struct and list columns become unusable text in CSV. Flatten them properly with DuckDB SQL, or keep the nesting by exporting JSON instead.
Convert a Large Parquet File to CSV Without Running Out of Memory
pandas dies converting a multi-GB Parquet file to CSV? Stream it instead — DuckDB COPY, pyarrow batches or Polars sink_csv write the CSV without loading it all.
How to Merge Multiple Parquet Files into One
Combine several Parquet files into a single dataset with DuckDB SQL in your browser — including files whose schemas don't quite match — or with pyarrow.
How to Open a Parquet File Without Spark or Python
Five ways to open and inspect a Parquet file — from zero-install browser tools to DuckDB one-liners — and when each one makes sense.
How to Open a Parquet File in Excel (Step by Step)
Excel cannot read Parquet natively. Convert the file to CSV in your browser — and trim oversized files with SQL first — to get it into a spreadsheet.
Why pandas Reads Parquet Integer Columns as Float (and the Fix)
An int column with nulls comes back from pd.read_parquet as float64, and large IDs lose digits. Find whether the file or the reader is at fault, and fix it.
Fix pandas read_parquet: Unable to Find a Usable Engine
pandas raises ImportError: Unable to find a usable engine when no Parquet reader is installed. What actually causes it, and how to fix it for good.
PARQUET_COLUMN_DATA_TYPE_MISMATCH: Find Which File Drifted
Spark or Databricks fails reading a Parquet dataset when one file's column type changed. Here is how to find the drifted file and column, and fix it.
Parquet Compression: Snappy vs Gzip vs Zstd, Which to Pick
Snappy, Gzip, and Zstd trade file size for CPU time differently. Benchmarks, defaults by engine, and how to check or change the codec in an existing Parquet file.
Why Your Parquet Timestamps Are Wrong: The INT96 Problem
Timestamps shifted by hours, wrong by centuries, or rejected outright? Your Parquet file probably uses legacy INT96. What it is and how to deal with it.
Parquet "Magic Bytes Not Found" — What It Means and How to Fix It
The PAR1 magic bytes error means your Parquet file is truncated or not Parquet at all. Here is how to diagnose it in seconds and where the file broke.
View Parquet Row Group Statistics (Min, Max, Null Count)
Read the per-row-group min/max and null counts stored in a Parquet footer with DuckDB's parquet_metadata(), and use them to explain why a filter is fast or slow.
Parquet to CSV: Control Nulls, Timestamps and Quoting
Your exported CSV has NaN, shifted timestamps or mangled accents. How to control null strings, date formats, delimiters and encoding when writing CSV.
Load a Parquet File into a DuckDB or SQLite Database
Turn Parquet files into a queryable .duckdb or .sqlite database file — with the right commands, the type traps, and when to skip the import entirely.
Parquet to JSON: Fix "Object Is Not JSON Serializable"
json.dumps fails on Timestamp, Decimal, bytes and int64 values read from Parquet. Why each type breaks, and three ways to get valid JSON out.
Convert Parquet to JSONL for LLM Fine-Tuning
Fine-tuning APIs expect JSONL, but datasets ship as Parquet. Filter, reshape and export training data as JSONL — in the browser or with one DuckDB command.
Convert Flat Parquet Rows to Nested JSON with DuckDB
Turn flat Parquet rows into nested JSON — orders with an items array, customer objects — using DuckDB's list() and struct syntax. Includes JSON array and JSONL output.
Parquet vs CSV: Size, Speed and When to Use Each
A practical comparison of Parquet and CSV — file size, query speed, type safety and tool support — with concrete guidance on choosing.
Print a Parquet File's Schema Without Installing Anything
See a Parquet file's columns, types and row count in seconds — in your browser, with no pyarrow, Spark or parquet-tools. Works for multi-GB files.
Query Parquet Files on S3 Without Downloading Them
Read Parquet straight from S3, GCS or any HTTPS URL with DuckDB httpfs — credentials, globs, partition pruning, and how to keep the bytes you pay for down.
Query Parquet Files with SQL — No Database Required
Run real SQL against local Parquet and CSV files using DuckDB in the browser: joins, aggregates and window functions with zero setup.
Read Parquet Key-Value Metadata (pandas, Spark, created_by)
Every Parquet footer stores key-value metadata: the writer, the pandas index, Spark's schema, custom tags. Read it with DuckDB SQL or pyarrow, and write your own.
How to Read a Partitioned Parquet Dataset Without Spark
Query a directory of Hive-partitioned Parquet files with DuckDB SQL, filter by partition columns, and understand how partition pruning works.
How to Rename, Drop or Reorder Columns in a Parquet File
Parquet files can't be edited in place. Rename, drop, reorder or retype columns by rewriting the file with one DuckDB statement or a few lines of pyarrow.
How to Split a Large Parquet File into Smaller Files
Split an oversized Parquet file by target size, by row count, or by a column value — with DuckDB SQL, pyarrow, or in the browser with no install.