diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 925f789b8..c991caeab 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -1,5 +1,7 @@ # Contributing to Pilosa +The workflow components of these instructions apply to all Pilosa repositories. + ## Reporting a bug If you have discovered a bug and don't see it in the [github issue tracker][5], [open a new issue][1]. @@ -20,7 +22,7 @@ If you want to help but you aren't sure where to start, check out our [github la ### Development Environment -- Ensure you have a recent version of [Go](https://golang.org/doc/install) installed. Pilosa generally supports the current and previous minor versions; check our [travis file](../.travis.yml) for the most up-to-date information. +- Ensure you have a recent version of [Go](https://golang.org/doc/install) installed. Pilosa generally supports the current and previous minor versions; check our [travis file](../master/.travis.yml) for the most up-to-date information. - Make sure `$GOPATH` environment variable points to your Go working directory and `$PATH` incudes `$GOPATH/bin`, as described [here](https://golang.org/doc/code.html#GOPATH). @@ -154,7 +156,7 @@ Additional commands are available in the `Makefile`. git checkout -b something-amazing ``` -- Commit your changes locally using `git add` and `git commit`. +- Commit your changes locally using `git add` and `git commit`. Please use [appropriate commit messages](https://chris.beams.io/posts/git-commit/). - Make sure that you've written tests for your new feature, and then run the tests: diff --git a/cmd/import.go b/cmd/import.go index 93fa92ee0..fe01738ad 100644 --- a/cmd/import.go +++ b/cmd/import.go @@ -61,7 +61,6 @@ omitted. If it is present then its format should be YYYY-MM-DDTHH:MM. flags.BoolVarP(&Importer.Sort, "sort", "", false, "Enables sorting before import.") flags.BoolVarP(&Importer.CreateSchema, "create", "e", false, "Create the schema if it does not exist before import.") flags.Var(&Importer.FrameOptions.TimeQuantum, "frame-time-quantum", "Time quantum for the frame") - flags.BoolVar(&Importer.FrameOptions.InverseEnabled, "frame-inverse-enabled", false, "Enable inverse frame") flags.BoolVar(&Importer.FrameOptions.RangeEnabled, "frame-range-enabled", false, "DEPRECATED - any frame can have fields. This option will be removed.") flags.StringVar(&Importer.FrameOptions.CacheType, "frame-cache-type", pilosa.CacheTypeRanked, "Cache type for the frame; valid values: none, lru, ranked") flags.Uint32Var(&Importer.FrameOptions.CacheSize, "frame-cache-size", 50000, "Cache size for the frame") diff --git a/cmd/root.go b/cmd/root.go index ec9d3be15..c3d5ca1bb 100644 --- a/cmd/root.go +++ b/cmd/root.go @@ -43,7 +43,7 @@ func NewRootCommand(stdin io.Reader, stdout, stderr io.Writer) *cobra.Command { This binary contains Pilosa itself, as well as common tools for administering pilosa, importing/exporting data, backing up, and more. Complete documentation is available -at https://www.pilosa.com/docs/ +at https://www.pilosa.com/docs/. ` + productName + ` Build Time: ` + pilosa.BuildTime + "\n", diff --git a/docs/README.md b/docs/README.md index 6a9f50b45..21fa269c0 100644 --- a/docs/README.md +++ b/docs/README.md @@ -1,5 +1,5 @@ Pilosa docs are maintained here, to stay in sync with the codebase. The format is [Blackfriday](https://github.com/russross/blackfriday) markdown, with some Hugo [front matter](https://gohugo.io/content-management/front-matter/). -Please visit [our website](https://www.pilosa.com/docs/) to view the docs complete with styles, diagrams, and comprehensive search. +Please visit [our website](https://www.pilosa.com/docs/) to view the docs complete with styles, diagrams, and comprehensive search. Internal links will only work on the website. Have you found a discrepancy, typo, or other problem? Please submit an [issue](https://github.com/pilosa/pilosa/issues/new) or a pull request! diff --git a/docs/administration.md b/docs/administration.md index e88d9842c..5a469c85e 100644 --- a/docs/administration.md +++ b/docs/administration.md @@ -20,23 +20,23 @@ Pilosa is a standalone, compiled Go application, so there is no need to worry ab #### Memory -Pilosa holds all row/column bitmap data in main memory. While this data is compressed more than a typical database, available memory is a primary concern. In a production environment, we recommend choosing hardware with a large amount of memory >= 64GB. Prefer a small number of hosts with lots of memory per host over a larger number with less memory each. Larger clusters tend to be less efficient overall due to increased inter-node communication. +Pilosa holds all row/column bitmap data in main memory. While this data is compressed more than a typical database, available memory is a primary concern. In a production environment, we recommend choosing hardware with a large amount of memory >= 64GB. Prefer a small number of hosts with lots of memory per host over a larger number with less memory each. Larger clusters tend to be less efficient overall due to increased inter-node communication. #### CPUs -Pilosa is a concurrent application written in Go and can take full advantage of multicore machines. The main unit of parallelism is the slice, so a single query will only use a number of cores up to the number of slices stored on that host. Multiple queries can still take advantage of multiple cores as well though, so tuning in this area is dependent on the expected workload. +Pilosa is a concurrent application written in Go and can take full advantage of multicore machines. The main unit of parallelism is the [slice](../data-model/#slice), so a single query will only use a number of cores up to the number of slices stored on that host. Multiple queries can still take advantage of multiple cores as well though, so tuning in this area is dependent on the expected workload. #### Disk -Even though the main dataset is in memory Pilosa does back up to disk frequently. We recommend SSDs--especially if you have a write heavy application. +Even though the main dataset is in memory Pilosa does back up to disk frequently. We recommend SSDs—especially if you have a write heavy application. #### Network -Pilosa is designed to be a distributed application, with data replication shared across the cluster. As such every write and read needs to communicate with several nodes. Therefore fast internode communication is essential. If using a service like AWS we recommend that all node exist in the same region and availability zone. The inherent latency of spreading a Pilosa cluster across physical regions it not usually worth the redundancy protection. Since Pilosa is designed to be an Indexing service there already should be a system of record, or ability to rebuild a Cluster quickly from backups. +Pilosa is designed to be a distributed application, with data replication shared across the cluster. As such every write and read needs to communicate with several nodes. Therefore fast internode communication is essential. If using a service like AWS we recommend that all node exist in the same region and availability zone. The inherent latency of spreading a Pilosa cluster across physical regions it not usually worth the redundancy protection. Since Pilosa is designed to be an indexing service there already should be a system of record, or ability to rebuild a cluster quickly from backups. #### Overview -While Pilosa does have some high system requirements it is not a best practice to set up a cluster with the fewest, largest machines available. You want an evenly distributed load across several nodes in a cluster to easily recover from a single node failure, and have the resource capacity to handle a missing node until it's repaired or replaced. Nor is it advisable to have many small machines. The internode network traffic will become a bottleneck. You can always add nodes later, but that does require some down time. +While Pilosa does have some high system requirements it is not a best practice to set up a cluster with the fewest, largest machines available. You want an evenly distributed load across several nodes in a cluster to easily recover from a single node failure, and have the resource capacity to handle a missing node until it's repaired or replaced. Nor is it advisable to have many small machines. The internode network traffic will become a bottleneck. You can always add nodes later, but that does require some down time. ### Open File Limits @@ -48,7 +48,7 @@ On Mac OS X, `ulimit` does not behave predictably. [This blog post](https://blog #### Importing -The import API expects a csv of rowID,columnID's. +The import API expects a csv of the format `Row,Column`. When importing large datasets remember it is much faster to pre sort the data by row ID and then by column ID in ascending order. You can use the `--sort` flag to do that. Also, avoid querying Pilosa until the import is complete, otherwise you will experience inconsistent results. @@ -58,7 +58,7 @@ pilosa import --sort -i project -f stargazer project-stargazer.csv ##### Importing Field Values -If you are using [BSI Range-Encoding](../data-model/#bsi-range-encoding) field values, you can import field values for a single frame and single field using `--field`. The CSV file should be in the format `ColumnID,Value`. +If you are using [BSI Range-Encoding](../data-model/#bsi-range-encoding) field values, you can import field values for a single frame and single field using `--field`. The CSV file should be in the format `Column,Value`. ``` pilosa import -i project -f stargazer --field star_count project-stargazer-counts.csv @@ -70,11 +70,18 @@ pilosa import -i project -f stargazer --field star_count project-stargazer-count #### Exporting -Exporting Data to csv can be performed on a live instance of Pilosa. You need to specify the Index, Frame, and View(default is standard). The API also expects the slice number, but the `pilosa export` sub command will export all slices within a Frame. The data will be in csv format rowID,columnID and sorted by columnID. -``` +Exporting data to csv can be performed on a live instance of Pilosa. You need to specify the index, frame, and view (default is standard). The API also expects the slice number, but the `pilosa export` sub command will export all slices within a Frame. The data will be in csv format `Row,Column` and sorted by column. +```request curl "http://localhost:10101/export?index=repository&frame=stargazer&slice=0&view=standard" \ --header "Accept: text/csv" ``` +```response +2,10 +2,30 +3,426 +4,2 +... +``` ### Versioning @@ -107,13 +114,13 @@ Pilosa v0.9 introduces a few compatibility changes that need to be addressed. **Configuration changes**: These changes need to occur before starting Pilosa v0.9: -1. Cluster-resize capability eliminates the `hosts` setting. Now, cluster membership is determined by `gossip`. This is only a factor if you are running Pilosa as a cluster. +1. Cluster-resize capability eliminates the `hosts` setting. Now, cluster membership is determined by gossip. This is only a factor if you are running Pilosa as a cluster. 2. Gossip-based cluster membership requires you to set a single cluster node as a [coordinator](../configuration/#cluster-coordinator). Make sure only a single node has the `cluster.coordinator` flag set. 3. `gossip.seed` has been renamed [`gossip.seeds`](../configuration/#gossip-seeds) and takes multiple items. It is recommended that at least two nodes are specified as gossip seeds. **Data directory changes**: These changes need to occur while the cluster is shut down, before starting Pilosa v0.9: -Pilosa v0.9 adds two new files to the data directory, an `.id` file and a `.topology` file. Due to the way Pilosa internally shards indices, upgrading a Pilosa cluster will result in data loss if an existing cluster is brought up without these files. New clusters will generate them automatically, but you may migrate an existing cluster by using a tool we called `topology-generator`: +Pilosa v0.9 adds two new files to the data directory, an `.id` file and a `.topology` file. Due to the way Pilosa internally shards indices, upgrading a Pilosa cluster will result in data loss if an existing cluster is brought up without these files. New clusters will generate them automatically, but you may migrate an existing cluster by using a tool we called [`topology-generator`](https://github.com/pilosa/upgrade-utils/tree/master/v0.9/topology-generator): 1. Observe the `cluster.hosts` configuration value in Pilosa v0.8. The ordering of the nodes in the config file is significant, as it determines shard (AKA slice) ownership. Pilosa v0.9 uses UUIDs for each node, and the ordering is alphabetical. 2. Install the `topology-generator`: `go get github.com/pilosa/upgrade-utils/v0.9/topology-generator`. @@ -204,9 +211,9 @@ curl localhost:10101/cluster/resize/set-coordinator \ ### Backup/restore -Pilosa continuously writes out the in-memory bitmap data to disk. This data is organized by Index->Frame->Views->Fragment->numbered slice files. These data files can be routinely backed up to restore nodes in a cluster. +Pilosa continuously writes out the in-memory bitmap data to disk. This data is organized by Index->Frame->Views->Fragment->numbered slice files. These data files can be routinely backed up to restore nodes in a cluster. -Depending on the size of your data you have two options. For a small dataset you can rely on the periodic anti-entropy sync process to replicate existing data back to this node. +Depending on the size of your data you have two options. For a small dataset you can rely on the periodic anti-entropy sync process to replicate existing data back to this node. For larger datasets and to make this process faster you could copy the relevant data files from the other nodes to the new one before startup. @@ -224,7 +231,7 @@ Note: This will only work when the replication factor is >= 2 - To accomplish this you will first need: - List of all indexes on your cluster - List of all frames in your indexes - - Max slice per index, listed in the /status endpoint + - Max slice per index, listed in the `/slices/max` endpoint - With this information you can query the `/fragment/nodes` endpoint and iterate over each slice - Using the list of slices owned by this node you will then need to manually: - setup a directory structure similar to the other nodes with a path for each Index/Frame @@ -239,19 +246,19 @@ Each Pilosa cluster is configured by default to share anonymous usage details wi - **Version:** Version string of the build. - **Host:** Host URI. -- **Cluster:** List of nodes in the Cluster. -- **NumNodes:** Number of nodes in the Cluster. -- **NumCPU:** Number of Cores per Node +- **Cluster:** List of nodes in the cluster. +- **NumNodes:** Number of nodes in the cluster. +- **NumCPU:** Number of cores per node - **BSIEnabled:** Bit Slice Index Frames in use. - **TimeQuantumEnabled:** Time Quantum Frames in use. -- **NumIndexes:** Number of Indexes in the Cluster. -- **NumFrames:** Number of Frames in the Cluster. -- **NumSlices:** Number of Slices in the Cluster. -- **NumViews:** Number of Views in the Cluster. +- **NumIndexes:** Number of indexes in the Cluster. +- **NumFrames:** Number of frames in the Cluster. +- **NumSlices:** Number of slices in the Cluster. +- **NumViews:** Number of views in the Cluster. - **OpenFiles:** Open file handle count. - **GoRoutines:** Go routine count. -You can opt-out of the Pilosa diagnostics reporting by setting either the command line configuration option `--metric.diagnostics=false`, use the `PILOSA_METRIC_DIAGNOSTICS` environment variable, or the TOML configuration file `[metric]` `diagnostics` option. +You can opt-out of the Pilosa diagnostics reporting by setting the command line configuration option `--metric.diagnostics=false`, the `PILOSA_METRIC_DIAGNOSTICS` environment variable, or the TOML configuration file `[metric]` `diagnostics` option. ### Metrics @@ -274,14 +281,14 @@ StatsD Tags adhere to the DataDog format (key:value), and we tag the following: #### Events We currently track the following events -- **Index:** The creation of a new Index. -- **Frame:** The creation of a new Frame. -- **MaxSlice:** The Creation of a new Slice. +- **Index:** The creation of a new index. +- **Frame:** The creation of a new frame. +- **MaxSlice:** The creation of a new Slice. - **SetBit:** Count of set bits. - **ClearBit:** Count of cleared bits. - **ImportBit:** During a bulk data import this represents the count of bits created. -- **SetRowAttrs:** Count of Attributes set per row. -- **SetColumnAttrs:** Count of Attributes set per collumn. +- **SetRowAttrs:** Count of attributes set per row. +- **SetColumnAttrs:** Count of attributes set per column. - **Bitmap:** Count of Bitmap queries. - **TopN:** Count of TopN queries. - **Union:** Count of Union queries. @@ -291,6 +298,6 @@ We currently track the following events - **Range:** Count of Range queries. - **Snapshot:** Event count when the snapshot process is triggered. - **BlockRepair:** Count of data blocks that were out of sync and repaired. -- **Garbage Collection:** Event count when Garbage Collection occurs. -- **Goroutines:** Number of running Goroutines. +- **GarbageCollection:** Event count when garbage collection occurs. +- **Goroutines:** Number of running goroutines. - **OpenFiles:** Number of open file handles associated with running Pilosa process ID. diff --git a/docs/api-reference.md b/docs/api-reference.md index c73083d3c..7841dff2e 100644 --- a/docs/api-reference.md +++ b/docs/api-reference.md @@ -63,7 +63,7 @@ curl -XDELETE localhost:10101/index/user `POST /index//query` -Sends a query to the Pilosa server with the given index. The request body is UTF-8 encoded text and response body is in JSON by default. +Sends a [query](../query-language/) to the Pilosa server with the given index. The request body is UTF-8 encoded text and response body is in JSON by default. ``` request curl localhost:10101/index/user/query \ @@ -153,6 +153,7 @@ curl -XDELETE localhost:10101/index/user/frame/language Creates a new field to store integer values in the given frame. The request payload is JSON, and it must contain the fields `type`, `min`, `max`. + * `type` (string): Field type, currently only "int" is supported. * `min` (int): Minimum value allowed for this field. * `max` (int): Maximum value allowed for this field. @@ -190,5 +191,9 @@ integration tests and not in a typical production workflow. Note that in a multi-node cluster, the cache is only recalculated on the node that receives the request. +``` request +curl -XGET localhost:10101/recalculate-caches +``` + Response: `204 No Content` diff --git a/docs/client-libraries.md b/docs/client-libraries.md index d5074a059..f445139b0 100644 --- a/docs/client-libraries.md +++ b/docs/client-libraries.md @@ -10,7 +10,7 @@ nav = [ ## Client Libraries -This section contains example code for client libraries in several languages. Please remember that when modeling your data in Pilosa, it is best to keep row and column ids sequential. It is not wise to use the output of a hash, or randomly distributed ids with Pilosa. +This section contains example code for client libraries in several languages. Please remember that when modeling your data in Pilosa, it is best to keep row and column ids sequential. It is best to avoid using the output of a hash or randomly distributed ids with Pilosa. ### Go @@ -55,12 +55,12 @@ func main() { // Which repositories did user 14 star: response, _ = client.Query(stargazer.Bitmap(14)) - fmt.Println("User 14 starred: ", response.Result().Bitmap.Bits) + fmt.Println("User 14 starred: ", response.Result().Bitmap().Bits) // What are the top 5 languages in the sample data? response, err = client.Query(language.TopN(5)) languageIDs := []uint64{} - for _, item := range response.Result().CountItems { + for _, item := range response.Result().CountItems() { languageIDs = append(languageIDs, item.ID) } fmt.Println("Top 5 languages: ", languageIDs) @@ -70,14 +70,14 @@ func main() { repository.Intersect( stargazer.Bitmap(14), stargazer.Bitmap(19))) - fmt.Println("Both user 14 and 19 starred:", response.Result().Bitmap.Bits) + fmt.Println("Both user 14 and 19 starred:", response.Result().Bitmap().Bits) // Which repositories were starred by user 14 or 19: response, _ = client.Query( repository.Union( stargazer.Bitmap(14), stargazer.Bitmap(19))) - fmt.Println("User 14 or 19 starred:", response.Result().Bitmap.Bits) + fmt.Println("User 14 or 19 starred:", response.Result().Bitmap().Bits) // Which repositories were starred by user 14 or 19 and were written in language 1: response, _ = client.Query( @@ -87,13 +87,22 @@ func main() { stargazer.Bitmap(19), ), language.Bitmap(1))) - fmt.Println("User 14 or 19 starred, written in language 1:", response.Result().Bitmap.Bits) + fmt.Println("User 14 or 19 starred, written in language 1:", response.Result().Bitmap().Bits) // Set user 99999 as a stargazer for repository 77777? client.Query(stargazer.SetBit(99999, 77777)) } ``` +Running the above program should produce output like this: +``` +User 14 starred: [1 2 3 362 368 391 396 409 416 430 436 450 454 460 461 464 466 469 470 483 484 486 490 491 503 504 514] +Top 5 languages: [5 1 4 9 13] +Both user 14 and 19 starred: [2 3 362 396 416 461 464 466 470 486] +User 14 or 19 starred: [1 2 3 361 362 368 376 377 378 382 386 388 391 396 398 400 409 411 412 416 426 428 430 435 436 450 452 453 454 456 460 461 464 465 466 469 470 483 484 486 487 489 490 491 500 503 504 505 512 514] +User 14 or 19 starred, written in language 1: [1 2 362 368 382 386 416 426 435 456 461 483 500 503 504 514] +``` + ### Python You can find the Python client library for Pilosa at our [Python Pilosa Repository](https://github.com/pilosa/python-pilosa). Check out its [README](https://github.com/pilosa/python-pilosa/blob/master/README.md) or [readthedocs](https://pilosa.readthedocs.io/en/latest/) for more information and installation instructions. @@ -167,6 +176,15 @@ print("User 14 or 19 starred, written in language 1:", mutually_starred) client.query(stargazer.setbit(99999, 77777)) ``` +Running the above program should produce output like this: +``` +('User 8 starred: ', [1L, 2L, 3L, 362L, 368L, 391L, 396L, 409L, 416L, 430L, 436L, 450L, 454L, 460L, 461L, 464L, 466L, 469L, 470L, 483L, 484L, 486L, 490L, 491L, 503L, 504L, 514L]) +('Top 5 languages: ', [5L, 1L, 4L, 9L, 13L]) +('Both user 14 and 19 starred:', [2L, 3L, 362L, 396L, 416L, 461L, 464L, 466L, 470L, 486L]) +('User 14 or 19 starred:', [1L, 2L, 3L, 361L, 362L, 368L, 376L, 377L, 378L, 382L, 386L, 388L, 391L, 396L, 398L, 400L, 409L, 411L, 412L, 416L, 426L, 428L, 430L, 435L, 436L, 450L, 452L, 453L, 454L, 456L, 460L, 461L, 464L, 465L, 466L, 469L, 470L, 483L, 484L, 486L, 487L, 489L, 490L, 491L, 500L, 503L, 504L, 505L, 512L, 514L]) +('User 14 or 19 starred, written in language 1:', [1L, 2L, 362L, 368L, 382L, 386L, 416L, 426L, 435L, 456L, 461L, 483L, 500L, 503L, 504L, 514L]) +``` + ### Java You can find the Java client library for Pilosa at our [Java Pilosa Repository](https://github.com/pilosa/java-pilosa). Check out its [README](https://github.com/pilosa/java-pilosa/blob/master/README.md) for more information and installation instructions. @@ -218,7 +236,7 @@ public class StarTrace { // What are the top 5 languages in the sample data: response = client.query(language.topN(5)); List top_languages = response.getResult().getCountItems(); - List languageIDs = new ArrayList<>(); + List languageIDs = new ArrayList(); for (CountResultItem item : top_languages) { languageIDs.add(item.getID()); } @@ -260,3 +278,12 @@ public class StarTrace { } } ``` + +Running the above program should produce output like this: +``` +User 14 starred: [1, 2, 3, 362, 368, 391, 396, 409, 416, 430, 436, 450, 454, 460, 461, 464, 466, 469, 470, 483, 484, 486, 490, 491, 503, 504, 514] +Top Languages: [5, 1, 4, 9, 13] +Both user 14 and 19 starred: [2, 3, 362, 396, 416, 461, 464, 466, 470, 486] +User 14 or 19 starred: [1, 2, 3, 361, 362, 368, 376, 377, 378, 382, 386, 388, 391, 396, 398, 400, 409, 411, 412, 416, 426, 428, 430, 435, 436, 450, 452, 453, 454, 456, 460, 461, 464, 465, 466, 469, 470, 483, 484, 486, 487, 489, 490, 491, 500, 503, 504, 505, 512, 514] +User 14 or 19 starred, written in language 1: [1, 2, 362, 368, 382, 386, 416, 426, 435, 456, 461, 483, 500, 503, 504, 514] +``` diff --git a/docs/configuration.md b/docs/configuration.md index 2d9d109b9..c6ed8fbd4 100644 --- a/docs/configuration.md +++ b/docs/configuration.md @@ -11,7 +11,7 @@ nav = [ ## Configuration -Pilosa can be configured through command line flags, environment variables, and/or a configuration file; configured options take precedence in that order. So if an option is specified in a command line flag, it will take precedence over the same option specified in the environment, which would take precedence over that same option specified in the configuration file. +Pilosa can be configured through command line flags, environment variables, and/or a configuration file; configured options take precedence in that order. So if an option is specified in a command line flag, it will take precedence over the same option specified in the environment, which will take precedence over that same option specified in the configuration file. All options are available in all three configuration types with the exception of the `--config` option which specifies the location of the config file, and therefore will not be used if it is present in the config file. @@ -117,19 +117,19 @@ The config file is in the [toml format](https://github.com/toml-lang/toml) and h #### Gossip Seeds -* Description: This specifies which internal host(s) should be used to initialize membership in the cluster. Typcially this can be the address of any available host in the cluster. For example, when starting a three-node cluster made up of `node0`, `node1`, and `node2`, the `gossip.seeds` for all three nodes can be configured to be the address of `node0`. Multiple seeds should be comma-separated in the flag and env forms. -* Flag: `--gossip.seeds="localhost:11101"` -* Env: `PILOSA_GOSSIP_SEEDS="localhost:11101"` +* Description: This specifies which internal host(s) should be used to initialize membership in the cluster. Typically this can be the address of any available host in the cluster. For example, when starting a three-node cluster made up of `node0`, `node1`, and `node2`, the `gossip.seeds` for all three nodes can be configured to be the address of `node0`. Multiple seeds should be comma-separated in the flag and env forms. +* Flag: `--gossip.seeds="localhost:11101,localhost:11110"` +* Env: `PILOSA_GOSSIP_SEEDS="localhost:11101,localhost:11110"` * Config: ```toml [gossip] - seeds = ["localhost:11101"] + seeds = ["localhost:11101", "localhost:11110"] ``` #### Gossip Key -* Description: Path to the file which contains the key to encrypt gossip communication. The contents of the file should be either 16, 24, or 32 bytes to select AES-128, AES-192, or AES-256 encryption. You can read from `/dev/random` device on UNIX-like systems to create the key file; e.g., `head -c 32 /dev/random > gossip.key32` creates a key file to use AES-256. +* Description: Path to the file which contains the key to encrypt gossip communication. The contents of the file should be either 16, 24, or 32 bytes to select AES-128, AES-192, or AES-256 encryption. You can read from `/dev/random` device on UNIX-like systems to create the key file; e.g., `head -c 32 /dev/random > gossip.key32` creates a key file to use AES-256. * Flag: `--gossip.key="/var/secret/gossip.key32"` * Env: `PILOSA_GOSSIP_KEY="/var/secret/gossip.key32"` * Config: @@ -204,7 +204,7 @@ The config file is in the [toml format](https://github.com/toml-lang/toml) and h * Description: Amount of time to collect cpu profiling data at startup if `profile.cpu` is set. * Flag: `--profile.cpu-time="30s"` -* Env: `PILOSA_PROFILE_CPU_TIME="30s" +* Env: `PILOSA_PROFILE_CPU_TIME="30s"` * Config: ```toml @@ -213,9 +213,9 @@ The config file is in the [toml format](https://github.com/toml-lang/toml) and h ``` #### Metric Service -* Description: Which stats service to use. Choose from [statsd, expvar, none]. +* Description: Which stats service to use for collecting [metrics](../administration/#metrics). Choose from [statsd, expvar, none]. * Flag: `--metric.service=statsd` -* Env: `PILOSA_METRIC_SERVICE=statsd' +* Env: `PILOSA_METRIC_SERVICE=statsd` * Config: ```toml @@ -226,7 +226,7 @@ The config file is in the [toml format](https://github.com/toml-lang/toml) and h #### Metric Host * Description: Address of the StatsD service host. * Flag: `--metric.host=localhost:8125` -* Env: `PILOSA_METRIC_HOST=localhost:8125' +* Env: `PILOSA_METRIC_HOST=localhost:8125` * Config: ```toml @@ -248,7 +248,7 @@ The config file is in the [toml format](https://github.com/toml-lang/toml) and h #### Metric Diagnostics -* Description: Enable reporting of limited usage statistics to Pilosa developers. To disable, set to false. +* Description: Enable [reporting](../administration/#diagnostics) of limited usage statistics to Pilosa developers. To disable, set to false. * Flag: `metric.diagnostics` * Env: `PILOSA_METRIC_DIAGNOSTICS` * Config: @@ -399,7 +399,7 @@ The same cluster which uses HTTPS instead of HTTP can be configured as follows. ### Example Cluster Configuration (HTTPS, same host) -You can run a cluster on the same host using the configuration above with a few changes. Gossip port and bind adress should be different for each node and a data directory should be accessed only by a single node. +You can run a cluster on the same host using the configuration above with a few changes. Gossip port and bind address should be different for each node and a data directory should be accessed only by a single node. #### Node 0 @@ -444,7 +444,7 @@ You can run a cluster on the same host using the configuration above with a few [gossip] port = 12002 - seed = "locahost:12000" + seed = "localhost:12000" key = "/home/pilosa/private/gossip.key32" [cluster] diff --git a/docs/data-model.md b/docs/data-model.md index 8375ae15f..9569b1f90 100644 --- a/docs/data-model.md +++ b/docs/data-model.md @@ -20,22 +20,22 @@ nav = [ The central component of Pilosa's data model is a boolean matrix. Each cell in the matrix is a single bit - if the bit is set, it indicates that a relationship exists between that particular row and column. -Rows and columns can represent anything (they could even represent the same set of things). Pilosa can associate arbitrary key/value pairs (referred to as attributes) to rows and columns, but queries and storage are optimized around the core matrix. +Rows and columns can represent anything (they could even represent the same set of things - a [bigraph](https://en.wikipedia.org/wiki/Bigraph)). Pilosa can associate arbitrary key/value pairs (referred to as attributes) to rows and columns, but queries and storage are optimized around the core matrix. -Pilosa lays out data first in rows, so queries which get all the set bits in one or many rows, or compute a combining operation on multiple rows such as Intersect or Union are the fastest. Pilosa also has the ability to categorize rows into different "frames" and quickly retrieve the top rows in a frame sorted by the number of bits set in each row. +Pilosa lays out data first in rows, so queries which get all the set bits in one or many rows, or compute a combining operation on multiple rows such as Intersect or Union are the fastest. Pilosa categorizes rows into different *frames* and quickly retrieves the top rows in a frame sorted by the number of bits set in each row. -Please note that Pilosa is most performant when row and column IDs are sequential starting from 0. You can deviate from this to some degree, but if you try to set a bit with column ID 2^63, bad things will start to happen. +Please note that Pilosa is most performant when row and column IDs are sequential starting from 0. You can deviate from this to some degree, but setting a bit with column ID 263 on a single-node cluster, for example, will not work well due to memory limitations. ![basic data model diagram](/img/docs/data-model.svg) *Basic data model diagram* ### Index -The purpose of the Index is to represent a data namespace. You cannot perform cross-index queries. Column-level attributes are global to the Index. +The purpose of the Index is to represent a data namespace. You cannot perform cross-index queries. ### Column -Column ids are sequential increasing integers and are common to all Frames within an Index. +Column ids are sequential increasing integers and are common to all Frames within an Index. A single column often corresponds to a record in a relational table, although other configurations are possible, and sometimes preferable. ### Row @@ -43,13 +43,51 @@ Row ids are sequential increasing integers namespaced to each Frame within an In ### Frame -Frames are used to segment and define different functional characteristics within your entire index. You can think of a Frame as a table-like data partition within your Index. +Frames are used to segment rows within an index, for example to define different functional groups. A frame might correspond to a single field in a relational table, where each row in a standard frame represents a single possible value of the field. Similarly, a frame with BSI values could represent all possible integer values of a field . -Row attributes are namespaced at the Frame level. +#### Relational Analogy + +The Pilosa index is a flexible structure; it can represent any sort of high-cardinality binary matrix. We have explored a number of modeling patterns in Pilosa use cases; one accessible example is a direct analogy to the relational model, summarized here. + +Entities: + + Relational | Pilosa +-------------|---------------------------------------------- + Database | N/A *(internal: Holder)* + Table | Index + Row | Column + Column | Frame + Value | Row + Value (int) | Field.Value (see [BSI](#bsi-range-encoding)) + +Simple queries: + + Relational | Pilosa +---------------------------------------------|------------------------------------ + `select ID from People where Name = 'Bob'` | `Bitmap(frame=Name, row=[Bob])` + `select ID from People where Age > 30` | `Range(frame=Default, Age > 30)` + `select ID from People where Member = true` | `Bitmap(frame=Member, row=[true])` + +In the relational model, joins are often necessary. Because Pilosa supports extremely high cardinality in both rows and columns, many types of joins are accomplished with basic Pilosa queries across multiple frames. For example, this SQL join: + +```sql +select AVG(p.Age) from People p +inner join PersonCar pc on pc.PersonID=p.ID +inner join Cars c on pc.CarID=c.ID +where c.Make = 'Ford' +``` + +can be accomplished with a Pilosa query like this (note that [Sum](../query-language/#sum) returns a json object containing both the sum and count, from which the average is easily computed): + +```pql +Sum(Bitmap(frame="Car-Make", row=[Ford]), frame=Default, field=Age) +``` + +This is one major component of Pilosa's ability to combine relationships from multiple data stores. #### Ranked -Ranked Frames maintain a sorted cache of column counts by Row ID (yielding the top rows by columns with a bit set in each). This cache facilitates the TopN query. The cache size defaults to 50,000 and can be set at Frame creation. +Ranked Frames maintain a sorted cache of column counts by Row ID (yielding the top rows by columns with a bit set in each). This cache facilitates the TopN query. The cache size defaults to 50,000 and can be set at Frame creation. ![ranked frame diagram](/img/docs/frame-ranked.svg) *Ranked frame diagram* @@ -67,17 +105,19 @@ Setting a time quantum on a frame creates extra views which allow Range queries ### Attribute -Attributes are arbitrary key/value pairs that can be associated to both rows or columns. This metadata is stored in a separate BoltDB data structure. +Attributes are arbitrary key/value pairs that can be associated with either rows or columns. This metadata is stored in a separate BoltDB data structure. + +Column-level attributes are common across an index. That is, each column attribute applies to all bits in the corresponding column, across all frames in an index. Row attributes apply to all bits in the corresponding row. ### Slice -Indexes are sharded into groups of columns called Slices - each Slice contains a fixed number of columns which is the SliceWidth. SliceWidth is a non-configurable constant set to 220. +Indexes are sharded into groups of columns called Slices. Each Slice contains a fixed number of columns, which is the SliceWidth. SliceWidth is a constant that can only be modified at compile time, and before ingesting data. The default value is 220. -Columns are sharded on a preset width, and each shard is referred to as a Slice. Slices are operated on in parallel, and they are evenly distributed across a cluster via a consistent hash algorithm. +Query operations run in parallel, and they are evenly distributed across a cluster via a consistent hash algorithm. ### View -Views represent the various data layouts within a Frame. The primary View is called Standard, and it contains the typical Row and Column data. Time-based Views are automatically generated for each time quantum. Views are internally managed by Pilosa, and never exposed directly via the API. This simplifies the functional interface from the physical data representation. +Views represent the various data layouts within a Frame. The primary View is called Standard, and it contains the typical Row and Column data. Time-based Views are automatically generated for each time quantum. Views are internally managed by Pilosa, and never exposed directly via the API. #### Standard @@ -85,7 +125,7 @@ The standard View contains the same Row/Column format as the input data. #### Time Quantums -If a Frame has a time quantum, then Views are generated for each of the defined time segments. For example, for a frame with a time quantum of `YMD`, the following `SetBit()` queries will result in the data described in the illustration below: +If a Frame has a time quantum, then Views are generated for each of the defined time segments. For example, for a frame with a time quantum of `YMD`, the following `SetBit()` queries will result in the data described in the diagram below: ``` SetBit(frame="A", row=8, col=3, timestamp="2017-05-18T00:00") @@ -97,12 +137,11 @@ SetBit(frame="A", row=8, col=3, timestamp="2017-05-19T00:00") #### BSI Range-Encoding -Bit-Sliced Indexing (BSI) is the storage method Pilosa uses to represent multi-bit integers in a bitmap index. Integers are stored as n-bit, range-encoded -bit-sliced indexes of base-2, along with an additional bitmap indicating "not null". This means that a 16-bit integer will require 17 bitmaps: one for each 0-bit of the 16 bit-slice components (the 1-bit does not need to be stored because with range-encoding the highest bit position is always 1) and one for the non-null bitmap. Pilosa can evaluate `Range`, `Min`, `Max`, and `Sum` queries on these BSI integers. +Bit-Sliced Indexing (BSI) is the storage method Pilosa uses to represent multi-bit integers in a bitmap index. Integers are stored as n-bit, range-encoded bit-sliced indexes of base-2, along with an additional bitmap indicating "not null". This means that a 16-bit integer will require 17 bitmaps: one for each 0-bit of the 16 bit-slice components (the 1-bit does not need to be stored because with range-encoding the highest bit position is always 1) and one for the non-null bitmap. Pilosa can evaluate `Range`, `Min`, `Max`, and `Sum` queries on these BSI integers. The result of a `Sum` query includes a count, which can be used to compute an average with no other overhead. -Internally Pilosa stores each BSI `field` as a `view` within a `frame`. The rows of the `view` are composed of the base-2 representation of the integer. Pilosa manages the base-2 offset and translation that efficiently packs the integer value within the minimum set of rows. +Internally Pilosa stores each BSI `field` as a `view` within a `frame`. The rows of the `view` contain the base-2 representations of the integer values. Pilosa manages the base-2 offset and translation that efficiently packs the integer value within the minimum set of rows. -For example, the following `SetFieldValue()` queries will result in the data described in the illustration below: +For example, the following `SetFieldValue()` queries will result in the data described in the diagram below: ``` SetFieldValue(col=1, frame="A", field0=1) diff --git a/docs/examples.md b/docs/examples.md index d7830a2a8..ab04b2bb5 100644 --- a/docs/examples.md +++ b/docs/examples.md @@ -65,13 +65,13 @@ Each column that we want to use must be mapped to a combination of frames and ro ##### 0 columns → 1 frame -cab_type: contains one row for each type of cab. Each column, representing one ride, has a bit set in exactly one row of this frame. The mapping is a simple enumeration, for example yellow=0, green=1, etc. The values of the bits in this frame are determined by the source of the data. That is, we're importing data from several disparate sources: NYC yellow taxi cabs, NYC green taxi cabs, and Uber cars. For each source, the single row to be set in the cab_type frame is constant. +**cab_type**: contains one row for each type of cab. Each column, representing one ride, has a bit set in exactly one row of this frame. The mapping is a simple enumeration, for example yellow=0, green=1, etc. The values of the bits in this frame are determined by the source of the data. That is, we're importing data from several disparate sources: NYC yellow taxi cabs, NYC green taxi cabs, and Uber cars. For each source, the single row to be set in the cab_type frame is constant. ##### 1 column → 1 frame The following three frames are mapped in a simple direct way from single columns of the original data. -dist_miles: each row represents rides of a certain distance. The mapping is simple: as an example, row 1 represents rides with a distance in the interval [0.5, 1.5]. That is, we round the floating point value of distance to an integer, and use that as the row ID directly. Generally, the mapping from a floating point value to a row ID could be arbitrary. The rounding mapping is concise to implement, which simplifies importing and analysis. As an added bonus, it's human-readable. We'll see this pattern used several times. +**dist_miles:** each row represents rides of a certain distance. The mapping is simple: as an example, row 1 represents rides with a distance in the interval [0.5, 1.5]. That is, we round the floating point value of distance to an integer, and use that as the row ID directly. Generally, the mapping from a floating point value to a row ID could be arbitrary. The rounding mapping is concise to implement, which simplifies importing and analysis. As an added bonus, it's human-readable. We'll see this pattern used several times. In PDK parlance, we define a Mapper, which is simply a function that returns integer row IDs. PDK has a number of predefined mappers that can be described with a few parameters. One of these is LinearFloatMapper, which applies a linear function to the input, and casts it to an integer, so the rounding is handled implicitly. In code: ```go @@ -123,7 +123,7 @@ These same objects are represented in the JSON definition file: } ``` -Here, we define a list of Mappers, each including a name, which we use to refer to the mapper later, in the list of BitMappers. We can also do this with Parsers, but a few simple Parsers that need no configuration are available by default. We also have a list of Fields, which is simply a map of field names to column indices. We use these names in the BitMapper definitions to keep things human-readable. +Here, we define a list of Mappers, each including a name, which we use to refer to the mapper later, in the list of BitMappers. We can also do this with Parsers, but a few simple Parsers that need no configuration are available by default. We also have a list of Fields, which is simply a map of field names (in the source data) to column indices (in Pilosa). We use these names in the BitMapper definitions to keep things human-readable. **total_amount_dollars:** Here we use the rounding mapping again, so each row represents rides with a total cost that rounds to the row's ID. The BitMapper definition is very similar to the previous one. @@ -163,7 +163,7 @@ durm := pdk.CustomMapper{ #### Import process -After designing this schema and mapping, we capture it in a JSON definition file that can be read by the PDK import tool. Running `pdk taxi` runs the import based on the information in this file. See [PDK](../pdk/) for more details on this process. +After designing this schema and mapping, we capture it in a JSON definition file that can be read by the PDK import tool. Running `pdk taxi` runs the import based on the information in this file. For more details, see the [PDK](../pdk/) section, or check out the [code](https://github.com/pilosa/pdk/tree/master/usecase/taxi) itself. #### Queries @@ -171,17 +171,23 @@ Now we can run some example queries. Count per cab type can be retrieved, sorted, with a single PQL call. -``` +```request TopN(frame=cab_type) ``` +```response +{"results":[[{"id":1,"count":1992943},{"id":0,"count":7057}]]} +``` High traffic location IDs can be retrieved with a similar call. These IDs correspond to latitude, longitude pairs, which can be recovered from the mapping that generates the IDs. -``` +```request TopN(frame=pickup_grid_id) ``` +```response +{"results":[[{"id":5060,"count":40620},{"id":4861,"count":38145},{"id":4962,"count":35268},...]]} +``` -Average of total_amount per passenger_count can be computed with some postprocessing. We use a small number of `TopN` calls to retrieve counts of rides by passenger_count, then use those counts to compute an average. +Average of `total_amount` per `passenger_count` can be computed with some postprocessing. We use a small number of `TopN` calls to retrieve counts of rides by passenger_count, then use those counts to compute an average. ```python queries = '' @@ -197,12 +203,16 @@ for pcount, topn in zip(pcounts, resp.json()['results']): average_amounts.append(float(wsum)/count) ``` +
+Note that the BSI-powered Sum query now provides an alternative approach to this kind of query. +
+ For more examples and details, see this [ipython notebook](https://github.com/pilosa/notebooks/blob/master/taxi-use-case.ipynb). ### Chemical similarity search
-This example uses the inverse frames feature, which is deprecated as of v0.9.0. This will soon be updated to reflect the current Pilosa API. +This example uses the inverse frames feature, which is deprecated as of v0.9.0. The example will soon be updated to reflect the current Pilosa API.
#### Overview diff --git a/docs/faq.md b/docs/faq.md index ef9f8bfc8..e1c0110c3 100644 --- a/docs/faq.md +++ b/docs/faq.md @@ -16,8 +16,10 @@ Pilosa is not a database in the traditional sense. While Pilosa does store data ### Where does Pilosa fit in my stack? -Pilosa sits on top of a data store or multiple data stores. -How is Pilosa different than Elasticsearch since they are both indexes? +Pilosa was designed to index the relationships in your data. Pilosa runs along with your existing stack, integrating with one or more backing data stores. Pilosa can connect through a stream platform like Kafka or application integration via [PDK](../pdk/). + +### How is Pilosa different from Elasticsearch since they are both indexes? + Elasticsearch is a search engine based on Lucene, and is therefore very good at indexing and searching large volumes of unstructured text. As it matures, Elasticsearch has continued to move into the analytics space, but its core data object is still the "document". Pilosa is specifically designed to index structured data and improve query speed. By representing data as the relationship between objects, and then storing those relationships in bitmaps, Pilosa can very efficiently search and compare many millions of data points while still maintaining a small memory footprint. ### How do I get my data into Pilosa? @@ -30,12 +32,11 @@ For the case where data is continually mutating, one would apply a parallel data ### What languages can I use with it? -There is currently client support for Go, Python, and Java. If you want to use Pilosa with a different language, you can access Pilosa via the Pilosa API. +There is currently [client support](../client-libraries/) for [Go](https://github.com/pilosa/go-pilosa), [Python](https://github.com/pilosa/python-pilosa), and [Java](https://github.com/pilosa/java-pilosa). If you want to use Pilosa with a different language, you can access Pilosa via the [Pilosa API](../api-reference/). ### Do you query Pilosa using SQL? -One can access Pilosa directly via the terminal using the Pilosa Query Language (PQL), but a typical implementation would use one of the Pilosa client libraries to integrate with an existing codebase. There is currently client support for Go, Python, and Java. - +One can access Pilosa directly via the terminal using the [Pilosa Query Language](../query-language/) (PQL), but a typical implementation would use one of the Pilosa client libraries to integrate with an existing codebase. There is currently client support for Go, Python, and Java. ### Replication on each node? diff --git a/docs/getting-started.md b/docs/getting-started.md index bee165643..c5383e81c 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -20,7 +20,7 @@ Any HTTP tool can be used to interact with the Pilosa server. The examples in th ### Starting Pilosa Follow the steps in the [Install](../installation/) document to install Pilosa. -Execute the following in a terminal to run Pilosa with the default configuration (Pilosa will be available at `localhost:10101`): +Execute the following in a terminal to run Pilosa with the default configuration (Pilosa will be available at [localhost:10101](http://localhost:10101)): ``` pilosa server ``` @@ -39,7 +39,7 @@ curl localhost:10101/status ### Sample Project -In order to better understand Pilosa's capabilities, we will create a sample project called "Star Trace" containing information about the top 1,000 most recently updated Github repositories which have "go" in their name. The Star Trace index will include data points such as programming language, tags, and stargazers—people who have starred a project. +In order to better understand Pilosa's capabilities, we will create a sample project called "Star Trace" containing information about 1,000 popular Github repositories which have "go" in their name. The Star Trace index will include data points such as programming language, tags, and stargazers—people who have starred a project. Although Pilosa doesn't keep the data in a tabular format, we still use the terms "columns" and "rows" when describing the data model. We put the primary objects in columns, and the properties of those objects in rows. For example, the Star Trace project will contain an index called "repository" which contains columns representing Github repositories, and rows representing properties like programming languages and tags. We can better organize the rows by grouping them into sets called Frames. So the "repository" index might have a "languages" frame as well as a "tags" frame. You can learn more about indexes and frames in the [Data Model](../data-model/) section of the documentation. @@ -218,14 +218,14 @@ Set user 99999 as a stargazer for repository 77777: ``` request curl localhost:10101/index/repository/query \ -X POST \ - -d 'SetBit(frame="stargazer", column=77777, row=99999)' + -d 'SetBit(frame="stargazer", col=77777, row=99999)' ``` ``` response {"results":[true]} ``` Please note that while user ID 99999 may not be sequential with the other column IDs, it is still a relatively low number. -Don't try to use arbitrary 64-bit integers as column or row IDs in Pilosa - this will lead to poor performance, out of memory errors, and more. +Don't try to use arbitrary 64-bit integers as column or row IDs in Pilosa - this will lead to problems such as poor performance and out of memory errors. diff --git a/docs/glossary.md b/docs/glossary.md index 9b883017d..fb1a46690 100644 --- a/docs/glossary.md +++ b/docs/glossary.md @@ -26,6 +26,8 @@ nav = [] [Frame](../data-model/#frame): Frames are used to group [rows](#row) into different categories. Row IDs are namespaced by frame such that the same row ID in a different frame refers to a different row. For [ranked](#topn) frames, rows are kept in sorted order within the frame. +[Gossip](https://en.wikipedia.org/wiki/Gossip_protocol): A protocol used by Pilosa for internal communication. + [Index](../data-model/#index): An Index is a top level container in Pilosa, analogous to a database in an RDBMS. Queries cannot operate across multiple indexes. [Jump Consistent Hash](https://arxiv.org/pdf/1406.2294v1.pdf): A fast, minimal memory, consistent hash algorithm that evenly distributes the workload even when the number of buckets changes. diff --git a/docs/installation.md b/docs/installation.md index a0632c34c..303490f6d 100644 --- a/docs/installation.md +++ b/docs/installation.md @@ -40,10 +40,10 @@ There are four ways to install Pilosa on MacOS: Use [Homebrew](https://brew.sh/) This binary contains Pilosa itself, as well as common tools for administering pilosa, importing/exporting data, backing up, and more. Complete documentation is available - at https://www.pilosa.com/docs/ + at https://www.pilosa.com/docs/. - Version: v0.4.0 - Build Time: 2017-06-08T19:44:21+0000 + Version: v0.9.0-64-gf053d9a5 + Build Time: 2018-05-14T22:14:01+0000 Usage: pilosa [command] @@ -60,10 +60,10 @@ There are four ways to install Pilosa on MacOS: Use [Homebrew](https://brew.sh/) inspect Get stats on a pilosa data file. restore Restore data to pilosa from a backup file. server Run Pilosa. - sort Sort import data for optimal import performance. Flags: -c, --config string Configuration file to read from. + -h, --help help for pilosa Use "pilosa [command] --help" for more information about a command. ``` @@ -101,16 +101,15 @@ There are four ways to install Pilosa on MacOS: Use [Homebrew](https://brew.sh/) This binary contains Pilosa itself, as well as common tools for administering pilosa, importing/exporting data, backing up, and more. Complete documentation is available - at https://www.pilosa.com/docs/ + at https://www.pilosa.com/docs/. - Version: v0.4.0 - Build Time: 2017-06-08T19:44:21+0000 + Version: v0.9.0-64-gf053d9a5 + Build Time: 2018-05-14T22:14:01+0000 Usage: pilosa [command] Available Commands: - backup Backup data from pilosa. bench Benchmark operations. check Do a consistency check on a pilosa data file. @@ -122,10 +121,10 @@ There are four ways to install Pilosa on MacOS: Use [Homebrew](https://brew.sh/) inspect Get stats on a pilosa data file. restore Restore data to pilosa from a backup file. server Run Pilosa. - sort Sort import data for optimal import performance. Flags: -c, --config string Configuration file to read from. + -h, --help help for pilosa Use "pilosa [command] --help" for more information about a command. ``` @@ -169,10 +168,10 @@ There are four ways to install Pilosa on MacOS: Use [Homebrew](https://brew.sh/) This binary contains Pilosa itself, as well as common tools for administering pilosa, importing/exporting data, backing up, and more. Complete documentation is available - at https://www.pilosa.com/docs/ + at https://www.pilosa.com/docs/. - Version: v0.4.0 - Build Time: 2017-06-08T19:44:21+0000 + Version: v0.9.0-64-gf053d9a5 + Build Time: 2018-05-14T22:14:01+0000 Usage: pilosa [command] @@ -189,11 +188,10 @@ There are four ways to install Pilosa on MacOS: Use [Homebrew](https://brew.sh/) inspect Get stats on a pilosa data file. restore Restore data to pilosa from a backup file. server Run Pilosa. - sort Sort import data for optimal import performance. - Flags: -c, --config string Configuration file to read from. + -h, --help help for pilosa Use "pilosa [command] --help" for more information about a command. ``` @@ -261,10 +259,10 @@ There are three ways to install Pilosa on Linux: download the binary (recommende This binary contains Pilosa itself, as well as common tools for administering pilosa, importing/exporting data, backing up, and more. Complete documentation is available - at https://www.pilosa.com/docs/ + at https://www.pilosa.com/docs/. - Version: v0.4.0 - Build Time: 2017-06-08T19:44:21+0000 + Version: v0.9.0-64-gf053d9a5 + Build Time: 2018-05-14T22:14:01+0000 Usage: pilosa [command] @@ -281,10 +279,10 @@ There are three ways to install Pilosa on Linux: download the binary (recommende inspect Get stats on a pilosa data file. restore Restore data to pilosa from a backup file. server Run Pilosa. - sort Sort import data for optimal import performance. Flags: -c, --config string Configuration file to read from. + -h, --help help for pilosa Use "pilosa [command] --help" for more information about a command. ``` @@ -328,10 +326,10 @@ There are three ways to install Pilosa on Linux: download the binary (recommende This binary contains Pilosa itself, as well as common tools for administering pilosa, importing/exporting data, backing up, and more. Complete documentation is available - at https://www.pilosa.com/docs/ + at https://www.pilosa.com/docs/. - Version: v0.4.0 - Build Time: 2017-06-08T19:44:21+0000 + Version: v0.9.0-64-gf053d9a5 + Build Time: 2018-05-14T22:14:01+0000 Usage: pilosa [command] @@ -348,10 +346,10 @@ There are three ways to install Pilosa on Linux: download the binary (recommende inspect Get stats on a pilosa data file. restore Restore data to pilosa from a backup file. server Run Pilosa. - sort Sort import data for optimal import performance. Flags: -c, --config string Configuration file to read from. + -h, --help help for pilosa Use "pilosa [command] --help" for more information about a command. ``` diff --git a/docs/introduction.md b/docs/introduction.md index d2ebb4662..305c9680e 100644 --- a/docs/introduction.md +++ b/docs/introduction.md @@ -12,7 +12,7 @@ Pilosa is an open source, distributed bitmap index. [//]: # (TODO insert a graphic here?) -It is designed primarly for speed and horizontal scalability. If you have data with billions of objects that can have millions of possible attributes, and you want to explore those relationships, Pilosa can help you. +It is designed primarily for speed and horizontal scalability. If you have data with billions of objects that can have millions of possible attributes, and you want to explore those relationships, Pilosa can help you. "What attributes are the most common?", "Which objects have these specific attributes?", "What groups of attributes often appear together?" Pilosa is designed to answer these types of queries in real time, suitable for use with high rate data streams, or to power a user interface. diff --git a/docs/pdk.md b/docs/pdk.md index 4d1e599ad..2ca4b6ee8 100644 --- a/docs/pdk.md +++ b/docs/pdk.md @@ -38,29 +38,29 @@ the record to arrive at that field. For example: This JSON object would result in the following Pilosa schema: -| Name | Field | Type | Size/Min | Max | -|----------------|-----------|--------|----------|------------| -| name | | ranked | 100000 | | -| favorite_foods | | ranked | 100000 | | -| default | | Ranked | 100000 | | -| | age | int | 0 | 2147483647 | -| location | | ranked | 1000 | | -| | latitude | int | 0 | 2147483647 | -| | longitude | int | 0 | 2147483647 | -| location-city | | ranked | 100000 | | -| location-state | | ranked | 100000 | | +| Name | Field | Type | Min | Max | Size | +|----------------|-----------|--------|-----|------------|--------| +| name | | ranked | | | 100000 | +| favorite_foods | | ranked | | | 100000 | +| default | | ranked | | | 100000 | +| | age | int | 0 | 2147483647 | | +| location | | ranked | | | 1000 | +| | latitude | int | 0 | 2147483647 | | +| | longitude | int | 0 | 2147483647 | | +| location-city | | ranked | | | 100000 | +| location-state | | ranked | | | 100000 | -All frames are created as ranked frames by default, and fields are created with -a minmum size of zero and a fixed maximum of 2147483647. Fields at the top level +All frames are created as ranked frames by default, with the cache size listed above. Fields are created with +a minimum size of zero and a fixed maximum of 2147483647. Fields at the top level are created in the default frame. Frames are a dash-separated concatenation of all key values in the path - you can see this with frames like location-city. -Most the options to `pdk kafka` are self-explanatory (kafka hosts, pilosa hosts, +Most of the options to `pdk kafka` are self-explanatory (kafka hosts, pilosa hosts, kafka topics, kafka group, etc.), but there are a few options that give some control over the way data is indexed, and ingestion performance. -* `--batch-size`: The batch size control how many set bits or values are batched up to be imported *per frame*. So for fields that have one value per record, you have to wait for `batch-size` records to come through before you'll see the data indexed in Pilosa. Fields like `favorite_foods` which can have multiple values could be indexed sooner. +* `--batch-size`: The batch size controls how many set bits or values are batched up to be imported *per frame*. So for fields that have one value per record, you have to wait for `batch-size` records to come through before you'll see the data indexed in Pilosa. Fields like `favorite_foods` which can have multiple values could be indexed sooner. * `--framer.collapse`: This is a list of strings which will be removed from the frame names created by dash-concatentating all names in the JSON path to a value. E.G. if "location" were listed in `framer.collapse`, then there would be frames named "city" and "state" rather than "location-city" and "location-state". * `--framer.ignore`: This allows you to skip indexing on any path containing these strings. If you have a field like email address or some other unique ID, you might not want to index it. * `--subject-path`: If nothing is passed for this option, then each record will be assigned a unique sequential column ID. If `subject-path` is specified, then the value at this path in the record will be mapped to a column ID. If the same value appears in another record, the same column ID will be used. diff --git a/docs/query-language.md b/docs/query-language.md index 7d9f22886..bbfe1cbb7 100644 --- a/docs/query-language.md +++ b/docs/query-language.md @@ -19,7 +19,7 @@ This section will provide a detailed reference and examples for the Pilosa Query {"results":[...]} ``` -There will be one item in the `results` array for each PQL query in the request. The type of each item in the array will depend on the type of query - each query in the reference below lists it's result type. +There will be one item in the `results` array for each PQL query in the request. The type of each item in the array will depend on the type of query - each query in the reference below lists its result type. #### Conventions @@ -64,7 +64,7 @@ SetBit(, , , **Description:** -`SetBit`, assigns a value of 1 to a bit in the binary matrix, thus associating the given row in the given frame with the given column. +`SetBit` assigns a value of 1 to a bit in the binary matrix, thus associating the given row in the given frame with the given column. **Result Type:** boolean @@ -75,21 +75,31 @@ A return value of `false` indicates that the bit was already set to 1 and nothin **Examples:** -``` +Set the bit at row 1, column 10: +```request SetBit(frame="stargazer", col=10, row=1) ``` - -This query illustrates setting a bit in the stargazer frame. User with id=1 has starred repository with id=10. - -SetBit also supports providing a timestamp. To write the date that a user starred a repository. +```response +{"results":[true]} ``` + +This sets a bit in the stargazer frame, representing that the user with id=1 has starred the repository with id=10. + +SetBit also supports providing a timestamp. To write the date that a user starred a repository: +```request SetBit(frame="stargazer", col=10, row=1, timestamp="2016-01-01T00:00") ``` - -Setting multiple bits in a single request: +```response +{"results":[true]} ``` + +Set multiple bits in a single request: +```request SetBit(frame="stargazer", col=10, row=1) SetBit(frame="stargazer", col=10, row=2) SetBit(frame="stargazer", col=20, row=1) SetBit(frame="stargazer", col=30, row=2) ``` +```response +{"results":[false,true,true,true]} +``` #### SetRowAttrs **Spec:** @@ -110,17 +120,23 @@ SetRowAttrs queries always return `null` upon success. **Examples:** -``` +Set attributes `username` and `active` on row 10: +```request SetRowAttrs(frame="stargazer", row=10, username="mrpi", active=true) ``` - -Set username value and active status for user 10. These are arbitrary key/value pairs which have no meaning to Pilosa. You can see the attributes you've set on a row with a [Bitmap](../query-language/#bitmap) query like so `Bitmap(frame="stargazer", stargazer_id=10)`. - +```response +{"results":[null]} ``` + +Set username value and active status for user 10. These are arbitrary key/value pairs which have no meaning to Pilosa. You can see the attributes you've set on a row with a [Bitmap](../query-language/#bitmap) query like so `Bitmap(frame="stargazer", row=10)`. + +Delete attribute `username` on row 10: +```request SetRowAttrs(frame="stargazer", row=10, username=null) ``` - -Delete username value for user 10. +```response +{"results":[null]} +``` #### SetColumnAttrs @@ -142,18 +158,42 @@ SetColumnAttrs queries always return `null` upon success. Setting a value of `nu **Examples:** -``` +Set attributes `stars`, `url`, and `active` on column 10: +```request SetColumnAttrs(col=10, stars=123, url="http://projects.pilosa.com/10", active=true) ``` - -Set url value and active status for project 10. These are arbitrary key/value pairs which have no meaning to Pilosa. You can see the attributes you've set on a column with a [Bitmap](../query-language/#bitmap) query like so `Bitmap(frame="stargazer", col=10)`. - +```response +{"results":[null]} ``` + +Set url value and active status for project 10. These are arbitrary key/value pairs which have no meaning to Pilosa. + +ColumnAttrs can be requested by adding the URL parameter `columnAttrs=true` to a query. For example: +```request +curl localhost:10101/index/repository/query?columnAttrs=true -XPOST -d 'Bitmap(frame="stargazer", row=1)Bitmap(frame="stargazer", row=2)' +``` +```response +{ + "results":[ + {"attrs":{},"bits":[10,20]}, + {"attrs":{},"bits":[10,30]} + ], + "columnAttrs":[ + {"id":10,"attrs":{"active":true,"stars":123,"url":"http://projects.pilosa.com/10"}}, + {"id":20,"attrs":{"active":false,"stars":456,"url":"http://projects.pilosa.com/30"}} + ] +} +``` + +In this example, ColumnAttrs have been set on columns 10 and 20, but not column 30. The relevant attributes are all returned in a single columnAttrs list. See the [query index](../api-reference/#query-index) section for more information. + +Delete the `url` attribute on column 10: +```request SetColumnAttrs(col=10, url=null) ``` - -Delete url value for repo 10. - +```response +{"results":[null]} +``` #### ClearBit @@ -166,7 +206,7 @@ ClearBit(, , , **Description:** -`ClearBit`, assigns a value of 0 to a bit in the binary matrix, thus disassociating the given row in the given frame from the given column. +`ClearBit` assigns a value of 0 to a bit in the binary matrix, thus disassociating the given row in the given frame from the given column. **Result Type:** boolean @@ -176,12 +216,15 @@ A return value of `false` indicates that the bit was already set to 0 and nothin **Examples:** -``` +Clear the bit at row 1 and column 10 in the stargazer frame: +```request ClearBit(frame="stargazer", col=10, row=1) ``` +```response +{"results":[true]} +``` -Remove relationship between the stargazer in row 1 and the repository in column 10 from the stargazer frame. - +This represents removing the relationship between the user with id=1 and the repository with id=10. #### SetFieldValue @@ -201,10 +244,17 @@ SetFieldValue returns `null` upon success. **Examples:** -Set the number of pull requests of repository 10. -``` +Set the field value `pullrequest` to the value 2, on column 10 in frame `stats`: +```request SetFieldValue(col=10, frame="stats", pullrequests=2) ``` +```response +{"results":[null]} +``` + +This represents setting the number of pull requests of repository 10 to 2. + +This example assumes the existence of the frame `stats` and the field `pullrequests`. See [frame creation](../api-reference/#create-frame) and [field creation](../api-reference/#create-field) for more information. ### Read Operations @@ -227,12 +277,13 @@ e.g. `{"attrs":{"username":"mrpi","active":true},"bits":[10, 20]}` **Examples:** -Query all repositories that user 1 has starred. -``` +Query all columns with a bit set in row 1 of the frame `stargazer` (repositories that are starred by user 1): +```request Bitmap(frame="stargazer", row=1) ``` - -Returns `{"attrs":{"username":"mrpi","active":true},"bits":[10, 20]}` +```response +{"attrs":{"username":"mrpi","active":true},"bits":[10, 20]} +``` * attrs are the attributes for user 1 * bits are the repositories which user 1 has starred. @@ -247,7 +298,7 @@ Union([BITMAP_CALL ...]) **Description:** -Union performs a logical OR on the results of each `BITMAP_CALL` query passed to it. +Union performs a logical OR on the results of all `BITMAP_CALL` queries passed to it. **Result Type:** object with attrs and bits @@ -255,12 +306,13 @@ attrs will always be empty **Examples:** -Query all repositories that are contributed by multiple users -``` +Query columns with a bit set in either of two rows (repositories that are starred by either of two users): +```request Union(Bitmap(frame="stargazer", stargazer_id=1), Bitmap(frame="stargazer", stargazer_id=2)) ``` - -Returns `{"attrs":{},"bits":[10, 20, 30]}`. +```response +{"attrs":{},"bits":[10, 20, 30]} +``` * bits are repositories that were starred by user 1 OR user 2 @@ -275,7 +327,7 @@ Intersect(, [BITMAP_CALL ...]) **Description:** -Intersect performs a logical AND on the results of each `BITMAP_CALL` query passed to it. +Intersect performs a logical AND on the results of all `BITMAP_CALL` queries passed to it. **Result Type:** object with attrs and bits @@ -283,13 +335,14 @@ attrs will always be empty **Examples:** -Query repositories which have been starred by two users. +Query columns with a bit set in both of two rows (repositories that are starred by both of two users): -``` +```request Intersect(Bitmap(frame="stargazer", row=1), Bitmap(frame="stargazer", row=2)) ``` - -Returns `{"attrs":{},"bits":[10]}`. +```response +{"attrs":{},"bits":[10]} +``` * bits are repositories that were starred by user 1 AND user 2 @@ -311,20 +364,23 @@ attrs will always be empty **Examples:** -Query repositories which have been starred by one user and not another. -``` +Query columns with a bit set in one row and not another (repositories that are starred by one user and not another): +```request Difference(Bitmap(frame="stargazer", row=1), Bitmap( frame="stargazer", row=2)) ``` - -Return `{"results":[{"attrs":{},"bits":[20]}]}` +```response +{"results":[{"attrs":{},"bits":[20]}]} +``` * bits are repositories that were starred by user 1 BUT NOT user 2 -``` +Query for the opposite difference: +```request Difference(Bitmap(frame="stargazer", row=2), Bitmap( frame="stargazer", row=1)) ``` - -Return `{"attrs":{},"bits":[30]}` +```response +{"attrs":{},"bits":[30]} +``` * Bits are repositories that were starred by user 2 BUT NOT user 1 @@ -346,13 +402,14 @@ attrs will always be empty **Examples:** -Query repositories which have been starred by two users. +Query columns with a bit set in exactly one of two rows (repositories that are starred by only one of two users): -``` +```request Xor(Bitmap(frame="stargazer", row=1), Bitmap(frame="stargazer", row=2)) ``` - -Returns `{"attrs":{},"bits":[30]}`. +```response +{"results":[{"attrs":{},"bits":[10,20,30]}]} +``` * bits are repositories that were starred by user 1 XOR user 2 (user 1 or user 2, but not both) @@ -371,12 +428,13 @@ Returns the number of set bits in the `BITMAP_CALL` passed in. **Examples:** -Query the number of repositories to which a user has contributed. -``` +Query the number of bits set in a row (the number of repositories a user has starred): +```request Count(Bitmap(frame="stargazer", row=1)) ``` - -Return `2` +```response +{"results":[1]} +``` * Result is the number of repositories that user 1 has starred. @@ -401,38 +459,56 @@ have the attribute specified by `field` with one of the values specified in **Caveats:** * Performing a TopN() query on a frame with cache type ranked will return the top bitmaps sorted by count in descending order. -* Frames with cache type lru will maintain an LRU (Least Recently Used) cache, thus a TopN() query on this type of frame will return bitmaps sorted in order of most recently set bit. -* The frame's cache size determines the number of sorted bitmaps to maintain in the cache for purposes of TopN() queries. There is a tradeoff between performance and accuracy; increasing the cache size will improve accuracy of results at the cost of performance. +* Frames with cache type lru will maintain an LRU (Least Recently Used replacement policy) cache, thus a TopN query on this type of frame will return bitmaps sorted in order of most recently set bit. +* The frame's cache size determines the number of sorted bitmaps to maintain in the cache for purposes of TopN queries. There is a tradeoff between performance and accuracy; increasing the cache size will improve accuracy of results at the cost of performance. * Once full, the cache will truncate the set of bitmaps according to the frame option CacheSize. Bitmaps that straddle the limit and have the same count will be truncated in no particular order. -* The TopN() query's attribute filter is applied to the existing sorted cache of bitmaps. Bitmaps that fall outside of the sorted cache range, even if they would normally pass the filter, are ignored. +* The TopN query's attribute filter is applied to the existing sorted cache of bitmaps. Bitmaps that fall outside of the sorted cache range, even if they would normally pass the filter, are ignored. + +See [frame creation](../api-reference/#create-frame) for more information about the cache. **Examples:** -``` +Basic TopN query: +```request TopN(frame="stargazer") ``` - -Returns `[{"key": 1, "count": 2}, {"key": 2, "count": 2}, {"key": 3, "count": 1}]` - -* key is a user ID -* count is amount of repositories -* Results are the number of repositories that each user starred in descending order for all users in the stargazer frame, for example user 1 starred two repositories, user 2 starred two repositories, user 3 starred one repository. - +```response +{"results":[[{"id":1240,"count":102},{"id":4734,"count":100},{"id":12709,"count":93},...]]} ``` + +* `id` is a row ID (user ID) +* `count` is a count of columns (repositories) +* Results are the number of bits set in the corresponding row (repositories that each user starred) in descending order for all rows (users) in the stargazer frame. For example user 1240 starred 102 repositories, user 4734 starred 100 repositories, user 12709 starred 93 repository. + +Limit the number of results: +```request TopN(frame="stargazer", n=2) ``` - -Returns `[{"key": 1, "count": 2}, {"key": 2, "count": 2}]` - -* Results are the top two users sorted by number of repositories they've starred in descending order. - +```response +{"results":[[{"id":1240,"count":102},{"id":4734,"count":100}]]} ``` + +* Results are the top two rows (users) sorted by number of bits set (repositories they've starred) in descending order. + +Filter based on an existing Bitmap: +```request TopN(Bitmap(frame="language", row=1), frame="stargazer", n=2) ``` +```response +{"results":[[{"id":1240,"count":35},{"id":7508,"count":32}]]} +``` -Returns `[{"key": 1, "count": 2}, {"key": 2, "count": 1}]` +* Results are the top two users (rows) sorted by the number of bits set in the intersection with row 1 of the language frame (repositories that they've starred which are written in language 1). -* Results are the top two users sorted by the number of repositories that they've starred which are written in language 1. +Filter based on attributes: +```request +TopN(frame="stargazer", n=2, field=active, filters=[true]) +``` +```response +{"results":[[{"id":10,"count":1},{"id":13,"count":1}]]} +``` + +* Results are the top two users (rows) which have the "active" attribute set to "true", sorted by the number of bits set (repositories that they've starred). #### Range Queries @@ -453,14 +529,17 @@ between the given `start` and `end` timestamps. **Examples:** -When you set timestamp using SetBit, you will able to query all repositories that a user has starred within a date range. -``` +Query all columns with a bit set in row 1 of a frame (repositories that a user has starred), within a date range: +```request Range(frame="stargazer", row=1, start="2010-01-01T00:00", end="2017-03-02T03:00") ``` +```response +{{"attrs":{},"bits":[10]} +``` -Returns `{{"attrs":{},"bits":[10]}` +This example assumes timestamps have been set on some bits. -* bits are repositories which were starred by user 1 from 2010-01-01 to 2017-03-02 +* bits are repositories which were starred by user 1 in the time range 2010-01-01 to 2017-03-02. #### Range (BSI) @@ -482,13 +561,14 @@ Returns bits that are true for the comparison operator. **Examples:** In our source data, commitactivity was counted over the last year. -The following greater-than `Range` query returns all repositories having more than 100 commits. +The following greater-than `Range` query returns all columns with a field value greater than 100 (repositories having more than 100 commits): -``` +```request Range(frame="stats", commitactivity > 100) ``` - -Returns `{{"attrs":{},"bits":[10]}` +```response +{{"attrs":{},"bits":[10]} +``` * bits are repositories which had at least 100 commits in the last year. @@ -504,9 +584,9 @@ BSI range queries support the following operators: `!=` | not-equal-to, NEQ | integer or `null` `><` | between, BETWEEN | [integer, integer] -The `BETWEEN` query specifies an interval with both bounds, using `><` operator, and a two-element list containing the lower and upper bounds of the interval: +The `BETWEEN` form specifies an interval with both bounds, using the `><` operator, and a two-element list containing the lower and upper bounds of the interval: -``` +```pql Range(frame="stats", commitactivity >< [100, 200]) ``` @@ -522,20 +602,21 @@ Min([BITMAP_CALL], , ) **Description:** -Returns the minimum value of all BSI integer values in the `field` in this `frame`. If the optional `Bitmap` call is supplied, only columns with set bits are considered, otherwise all collumns are considered. +Returns the minimum value of all BSI integer values in the `field` in this `frame`. If the optional `Bitmap` call is supplied, only columns with set bits are considered, otherwise all columns are considered. **Result Type:** object with the min and count of columns containing the min value. **Examples:** -Query the size of all repositories. -``` +Query the minimum value of all fields in a frame (minimum size of all repositories): +```request Min(frame="stats", field="diskusage") ``` +```response +{"value":4,"count":2} +``` -Return `{"min":4,"count":2}` - -* Result is the smallest repository in kilobytes, plus the number of repositories of that size. +* Result is the smallest value (repository size in kilobytes, here), plus the count of columns with that value. #### Max @@ -553,14 +634,15 @@ Returns the maximum value of all BSI integer values in the `field` in this `fram **Examples:** -Query the size of all repositories. -``` +Query the maximum value of all fields in a frame (maximum size of all repositories): +```request Max(frame="stats", field="diskusage") ``` +```response +{"value":88,"count":13} +``` -Return `{"max":88,"count":13}` - -* Result is the largest repository in kilobytes, plus the number of repositories of that size. +* Result is the largest value (repository size in kilobytes, here), plus the count of columns with that value. #### Sum @@ -572,17 +654,18 @@ Sum([BITMAP_CALL], , ) **Description:** -Returns the count and computed sum of all BSI integer values in the `field` in this `frame`. If the optional `Bitmap` call is supplied, columns with set bits are summed, otherwise the sum is across all columns. +Returns the count and computed sum of all BSI integer values in the `field` and `frame`. If the optional `Bitmap` call is supplied, columns with set bits are summed, otherwise the sum is across all columns. **Result Type:** object with the computed sum and count of the bitmap field. **Examples:** Query the size of all repositories. -``` +```request Sum(frame="stats", field="diskusage") ``` +```response +{"value":10,"count":3} +``` -Return `{"sum":10,"count":3}` - -* Result is the size of all repositories in kilobytes, plus the number of repositories. +* Result is the sum of all values (total size of all repositories in kilobytes, here), plus the count of columns. diff --git a/docs/tutorials.md b/docs/tutorials.md index 39e345dbd..432c7466e 100644 --- a/docs/tutorials.md +++ b/docs/tutorials.md @@ -46,7 +46,7 @@ mkdir $HOME/pilosa-tls-tutorial && cd $_ Securing a Pilosa cluster consists of securing the communication between nodes using TLS and Gossip encryption. [Pilosa Enterprise](https://www.pilosa.com/enterprise/) additionally supports authentication and other security features, but those are not covered in this tutorial. -The first step is acquiring an SSL certificate. You can buy a commercial certificate or retrieve a Let's Encrypt certificiate but we will be using a self signed certificate for practical reasons. Using self-signed certificates is not recommended in production, since it makes man in the middle attacks easy. +The first step is acquiring an SSL certificate. You can buy a commercial certificate or retrieve a Let's Encrypt certificate but we will be using a self signed certificate for practical reasons. Using self-signed certificates is not recommended in production, since it makes man in the middle attacks easy. The following command creates a 2048bit self-signed wildcard certificate for `*.pilosa.local` which expires 10 years later. @@ -178,7 +178,7 @@ cd $HOME/pilosa-tls-tutorial pilosa server -c node3.config.toml ``` -Let's ensure that all three Pilosa servers are runnning and they are connected: +Let's ensure that all three Pilosa servers are running and they are connected: ``` curl -k --ipv4 https://01.pilosa.local:10501/status ``` diff --git a/docs/webui.md b/docs/webui.md index 0f14934b4..07acd2ab4 100644 --- a/docs/webui.md +++ b/docs/webui.md @@ -15,7 +15,7 @@ This can be used for constructing queries and viewing the cluster status. ### Console -The [Console view](http://localhost:10101/#console) allows you to enter [PQL](../query-language/) queries and run them against your locally running server. First you must select an Index with the Select index dropdown. +The [Console view](http://localhost:10101/#console) allows you to enter [PQL](../query-language/) queries and run them against your locally running server. First you must select an Index with the Select index dropdown. Each query's result will be displayed in the Output section along with the query time. @@ -39,4 +39,4 @@ Frame creation also supports options like `timeQuantum`. When creating a new fra ### Cluster Admin -Use the [Cluster Admin tab](http://localhost:10101/#admin) to view the current status of your cluster. This contains information on each node in the cluster, plus the list of Indexes and Frames. +Use the [Cluster Admin tab](http://localhost:10101/#admin) to view the current status of your cluster. This contains information on each node in the cluster, plus the list of Indexes and Frames.