Setup
Let’s first create a virtual environment:ipython to run the commands in the rest of the guide, which you can launch by running:
Exploring Parquet metadata
We’re going to explore a Parquet file from the Amazon reviews dataset. But first, let’s installchDB:
ParquetMetadata input format to have it return Parquet metadata rather than the content of the file.
Let’s use the DESCRIBE clause to see the fields returned when we use this format:
columns and row_groups both contain arrays of tuples containing many properties, so we’ll exclude those for now.
Querying Parquet files
Next, let’s query the contents of the file. We can do this by adjusting the above query to removeParquetMetadata and then, say, compute the most popular star_rating across all reviews: