- insert the data directly into ClickHouse Cloud from S3 or GCS
- download prepared partitions
- Alternatively you can query the full dataset in our demo environment at sql.clickhouse.com.
The example queries below were executed on a Production instance of ClickHouse Cloud. For more information see
“Playground specifications”.
Create the table trips
Start by creating a table for the taxi rides:Load the data directly from object storage
Users’ can grab a small subset of the data (3 million rows) for getting familiar with it. The data is in TSV files in object storage, which is easily streamed into ClickHouse Cloud using thes3 table function.
The same data is stored in both S3 and GCS; choose either tab.
- S3
- GCS
The following command streams three files from an S3 bucket into the
trips_small table (the {0..2} syntax is a wildcard for the values 0, 1, and 2):Sample queries
The following queries are executed on the sample described above. You can run the sample queries on the full dataset in sql.clickhouse.com, modifying the queries below to use the tablenyc_taxi.trips.
Let’s see how many rows were inserted:
Each TSV file has about 1M rows, and the three files have 3,000,317 rows. Let’s look at a few rows:
Notice there are columns for the pickup and dropoff dates, geo coordinates, fare details, New York neighborhoods, and more.
Let’s run a few queries. This query shows us the top 10 neighborhoods that have the most frequent pickups:
This query shows the average fare based on the number of passengers:
runnable
runnable
Download of prepared partitions
The following steps provide information about the original dataset, and a method for loading prepared partitions into a self-managed ClickHouse server environment.
If you will run the queries described below, you have to use the full table name,
datasets.trips_mergetree.