diff --git a/docs/getting-started.md b/docs/getting-started.md index 2e0104591..73971895e 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -11,7 +11,7 @@ nav = [ ## Getting Started Pilosa supports an HTTP interface which uses JSON by default. -Any HTTP tool can be used to interact with the Pilosa server. The examples in this documentation will use [curl](https://curl.haxx.se/) which is available by default on many UNIX-like systems including Linux and MacOS. Windows users can download curl [here](https://curl.haxx.se/download.html). +Any HTTP tool can be used to interact with the Pilosa server. The examples in this documentation will use curl which is available by default on many UNIX-like systems including Linux and MacOS. However, the best way to interface with the Pilosa server is through one of our three client libraries. Pilosa currently supports go, java, and python.
Note that Pilosa server requires a high limit for open files. Check the documentation of your system to see how to increase it in case you hit that limit. See Open File Limits for more details.
@@ -44,31 +44,196 @@ In order to better understand Pilosa's capabilities, we will create a sample pro Although Pilosa doesn't keep the data in a tabular format, we still use the terms "columns" and "rows" when describing the data model. We put the primary objects in columns, and the properties of those objects in rows. For example, the Star Trace project will contain an index called "repository" which contains columns representing Github repositories, and rows representing properties like programming languages and tags. We can better organize the rows by grouping them into sets called Fields. So the "repository" index might have a "languages" field as well as a "tags" field. You can learn more about indexes and fields in the [Data Model](../data-model/) section of the documentation. -#### Create the Environment +Note: +If at any time you want to verify the data structure, you can request the schema as follows: -While we can create indexes and query directly in the terminal, it is more conventional to do so in a client library. Pilosa supports Go, Java, and Python, though you will have to install the library for compatibility. +``` request +curl localhost:10101/schema +``` +``` response +{"indexes":null} +``` -For Go users, open a terminal (one other than the one running pilosa) and download the library in your `GOPATH` using: +#### Using HTTP + +Note: This is not the recommended way to interact with Pilosa, but it is the fastest way to see the efficiency of Pilosa. + +##### Creating the Schema + +Before we can import data or run queries, we need to create our indexes and the fields within them. Let's create the repository index first: +``` request +curl localhost:10101/index/repository -X POST +``` +``` response +{"success":true} +``` +The index name must be 64 characters or less, start with a letter, and consist only of lowercase alphanumeric characters or `_-`. + +Let's create the `stargazer` field which has user IDs of stargazers as its rows: +``` request +curl localhost:10101/index/repository/field/stargazer \ + -X POST \ + -d '{"options": {"type": "time", "timeQuantum": "YMD"}}' +``` +``` response +{"success":true} +``` + +Since our data contains time stamps which represent the time users starred repos, we set the field type to `time`. Time quantum is the resolution of the time we want to use, and we set it to `YMD` (year, month, day) for `stargazer`. + +Next up is the `language` field, which will contain IDs for programming languages: +``` request +curl localhost:10101/index/repository/field/language \ + -X POST +``` +``` response +{"success":true} +``` + +The `language` is a `set` field, but since the default field type is `set`, we didn't specify it in field options. + +##### Import Data From CSV Files + +Download the `stargazer.csv` and `language.csv` files here: + +``` +curl -O https://raw.githubusercontent.com/pilosa/getting-started/master/stargazer.csv +curl -O https://raw.githubusercontent.com/pilosa/getting-started/master/language.csv +``` + +Run the following commands to import the data into Pilosa: + +``` +pilosa import -i repository -f stargazer stargazer.csv +pilosa import -i repository -f language language.csv +``` + +Note that both the user IDs and the repository IDs were remapped to sequential integers in the data files, they don't correspond to actual Github IDs anymore. You can check out [languages.txt](https://github.com/pilosa/getting-started/blob/master/languages.txt) to see the mapping for languages. + +##### Make Some Queries + +Which repositories did user 14 star: +``` request +curl localhost:10101/index/repository/query \ + -X POST \ + -d 'Row(stargazer=14)' +``` +``` response +{ + "results":[ + { + "attrs":{}, + "columns":[1,2,3,362,368,391,396,409,416,430,436,450,454,460,461,464,466,469,470,483,484,486,490,491,503,504,514] + } + ] +} +``` + +What are the top 5 languages in the sample data: +``` request +curl localhost:10101/index/repository/query \ + -X POST \ + -d 'TopN(language, n=5)' +``` +``` response +{ + "results":[ + [ + {"id":5,"count":119}, + {"id":1,"count":50}, + {"id":4,"count":48}, + {"id":9,"count":31}, + {"id":13,"count":25} + ] + ] +} +``` + +Which repositories were starred by user 14 and 19: +``` request +curl localhost:10101/index/repository/query \ + -X POST \ + -d 'Intersect( + Row(stargazer=14), + Row(stargazer=19) + )' +``` +``` response +{ + "results":[ + { + "attrs":{}, + "columns":[2,3,362,396,416,461,464,466,470,486] + } + ] +} +``` + +Which repositories were starred by user 14 or 19: +``` request +curl localhost:10101/index/repository/query \ + -X POST \ + -d 'Union( + Row(stargazer=14), + Row(stargazer=19) + )' +``` +``` response +{ + "results":[ + { + "attrs":{}, + "columns":[1,2,3,361,362,368,376,377,378,382,386,388,391,396,398,400,409,411,412,416,426,428,430,435,436,450,452,453,454,456,460,461,464,465,466,469,470,483,484,486,487,489,490,491,500,503,504,505,512,514] + } + ] +} +``` + +Which repositories were starred by user 14 and 19 and also were written in language 1: +``` request +curl localhost:10101/index/repository/query \ + -X POST \ + -d 'Intersect( + Row(stargazer=14), + Row(stargazer=19), + Row(language=1) + )' +``` +``` response +{ + "results":[ + { + "attrs":{}, + "columns":[2,362,416,461] + } + ] +} +``` + +Set user 99999 as a stargazer for repository 77777: +``` request +curl localhost:10101/index/repository/query \ + -X POST \ + -d 'Set(77777, stargazer=99999)' +``` +``` response +{"results":[true]} +``` + +Please note that while user ID 99999 may not be sequential with the other column IDs, it is still a relatively low number. +Don't try to use arbitrary 64-bit integers as column or row IDs in Pilosa - this will lead to problems such as poor performance and out of memory errors. + +#### Using Go + +Pilosa requires Go 1.12 or higher. It is also recommended that you have a code editor downloaded. + +##### Create the Environment + +In order to communicate with Pilosa through your go code, you must have a "translator," which is go-pilosa. To install go-pilsa, open a terminal (one other than the one running pilosa) and download the library in your `GOPATH` using: ``` go get github.com/pilosa/go-pilosa ``` -For Java users, add the following dependency in your `pom.xml`: -``` -Java and Python support will be uploaded shortly. -
Java and Python support will be uploaded shortly. -
If you are using a Docker container for Pilosa (with name `pilosa`), you should instead copy the `*.csv` file into the container and then import them: -``` -docker cp stargazer.csv pilosa:/stargazer.csv -docker exec -it pilosa /pilosa import -i repository -f stargazer /stargazer.csv -docker cp language.csv pilosa:/language.csv -docker exec -it pilosa /pilosa import -i repository -f language /language.csv -``` -
Java and Python support will be uploaded shortly.