From 73904db391e548b72cb9277eb2708b4cc94817c4 Mon Sep 17 00:00:00 2001 From: Alan Bernstein Date: Wed, 21 Feb 2018 15:42:35 -0600 Subject: [PATCH 1/6] Fix indentation --- docs/administration.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/administration.md b/docs/administration.md index b77d05262..97ee688f5 100644 --- a/docs/administration.md +++ b/docs/administration.md @@ -124,7 +124,7 @@ Note: This will only work when the replication factor is >= 2 - Restart the cluster - Wait for the 1st sync (10 minutes) to validate Index connections -#### Diagnostics +### Diagnostics Each Pilosa cluster is configured by default to share anonymous usage details with Pilosa Corp. These metrics allow us to understand how Pilosa is used by the community and improve the technology to suit your needs. Diagnostics are sent to Pilosa every hour. Each of the metrics are detailed below as well as opt-out instructions. @@ -145,7 +145,7 @@ Each Pilosa cluster is configured by default to share anonymous usage details wi You can opt-out of the Pilosa diagnostics reporting by setting either the command line configuration option `--metric.diagnostics=false`, use the `PILOSA_METRIC_DIAGNOSTICS` environment variable, or the TOML configuration file `[metric]` `diagnostics` option. -#### Metrics +### Metrics Pilosa can be configured to emit metrics pertaining to its internal processes in one of two formats: Expvar or StatsD. Metric recording is disabled by default. The metrics configuration options are: @@ -154,7 +154,7 @@ The metrics configuration options are: - [Poll Interval](../configuration#metrics-poll-interval): specify polling interval for runtime metrics - [Service](../configuration#metrics-service): declare type StatsD or Expvar -##### Tags +#### Tags StatsD Tags adhere to the DataDog format (key:value), and we tag the following: - NodeID @@ -163,7 +163,7 @@ StatsD Tags adhere to the DataDog format (key:value), and we tag the following: - View - Slice -##### Events +#### Events We currently track the following events - **Index:** The creation of a new Index. From 59d41a9cb3b3b938aef8ce18f8adc666fca49dc5 Mon Sep 17 00:00:00 2001 From: Alan Bernstein Date: Wed, 21 Feb 2018 15:42:46 -0600 Subject: [PATCH 2/6] Add Xor to docs --- docs/query-language.md | 24 +++++++++++++++++++++++- 1 file changed, 23 insertions(+), 1 deletion(-) diff --git a/docs/query-language.md b/docs/query-language.md index 0e02351ab..b1315d293 100644 --- a/docs/query-language.md +++ b/docs/query-language.md @@ -52,7 +52,7 @@ curl localhost:10101/index/repository/query \ * `UINT` An unsigned integer (e.g. 42839) * `ATTR_NAME` Must be a valid identifier `[A-Za-z][A-Za-z0-9._-]*` * `ATTR_VALUE` Can be a string, float, integer, or bool. -* `BITMAP_CALL` Any query which returns a bitmap, such as `Bitmap`, `Union`, `Difference`, `Intersect`, `Range` +* `BITMAP_CALL` Any query which returns a bitmap, such as `Bitmap`, `Union`, `Difference`, `Xor`, `Intersect`, `Range` * `[]ATTR_VALUE` Denotes an array of `ATTR_VALUE`s. (e.g. `["a", "b", "c"]`) ### Write Operations @@ -308,6 +308,28 @@ Return `{"attrs":{},"bits":[30]}` * Bits are repositories that were starred by user 2 BUT NOT user 1 +#### Xor + +**Description:** + +Xor performs a logical XOR on the results of each `BITMAP_CALL` query passed to it. + +**Result Type:** object with attrs and bits + +attrs will always be empty + +**Examples:** + +Query repositories which have been starred by two users. + +``` +Xor(Bitmap(frame="stargazer", rowID=1), Bitmap(frame="stargazer", rowID=2)) +``` + +Returns `{"attrs":{},"bits":[30]}`. + +* bits are repositories that were starred by user 1 XOR user 2 (user 1 or user 2, but not both) + #### Count **Spec:** From a433a5ebfc3aea405b97170969adf8b7d53c3319 Mon Sep 17 00:00:00 2001 From: Matthew Jaffee Date: Wed, 21 Feb 2018 16:42:35 -0600 Subject: [PATCH 3/6] remove column label repo_id from query language docs column labels are deprecated and the docs' usage of columnID vs repo_id was inconsisent. --- docs/query-language.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/docs/query-language.md b/docs/query-language.md index 0e02351ab..ee3d49a96 100644 --- a/docs/query-language.md +++ b/docs/query-language.md @@ -80,14 +80,14 @@ A return value of `false` indicates that the bit was already set to 1 and nothin **Examples:** ``` -SetBit(frame="stargazer", repo_id=10, rowID=1) +SetBit(frame="stargazer", columnID=10, rowID=1) ``` This query illustrates setting a bit in the stargazer frame. User with id=1 has starred repository with id=10. SetBit also supports providing a timestamp. To write the date that a user starred a repository. ``` -SetBit(frame="stargazer", repo_id=10, rowID=1, timestamp="2016-01-01T00:00") +SetBit(frame="stargazer", columnID=10, rowID=1, timestamp="2016-01-01T00:00") ``` Setting multiple bits in a single request: @@ -150,7 +150,7 @@ SetColumnAttrs queries always return `null` upon success. Setting a value of `nu SetColumnAttrs(columnID=10, stars=123, url="http://projects.pilosa.com/10", active=true) ``` -Set url value and active status for project 10. These are arbitrary key/value pairs which have no meaning to Pilosa. You can see the attributes you've set on a column with a [Bitmap]({{< ref "query-language.md#bitmap" >}}) query like so `Bitmap(frame="stargazer", repo_id=10)`. +Set url value and active status for project 10. These are arbitrary key/value pairs which have no meaning to Pilosa. You can see the attributes you've set on a column with a [Bitmap]({{< ref "query-language.md#bitmap" >}}) query like so `Bitmap(frame="stargazer", columnID=10)`. ``` SetColumnAttrs(columnID=10, url=null) @@ -184,7 +184,7 @@ A return value of `false` indicates that the bit was already set to 0 and nothin ClearBit(frame="stargazer", columnID=10, rowID=1) ``` -Remove relationship between stargazer_id 1 and repo_id 10 from the stargazer frame. +Remove relationship between the stargazer in row 1 and the repository in column 10 from the stargazer frame. ### Read Operations From a20235f812d5e6b0d0cc85599c404284b97c6f02 Mon Sep 17 00:00:00 2001 From: Alan Bernstein Date: Mon, 26 Feb 2018 15:26:28 -0600 Subject: [PATCH 4/6] Add Xor spec --- docs/query-language.md | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/docs/query-language.md b/docs/query-language.md index b1315d293..365d7ded6 100644 --- a/docs/query-language.md +++ b/docs/query-language.md @@ -310,6 +310,12 @@ Return `{"attrs":{},"bits":[30]}` #### Xor +**Spec:** + +``` +Xor(, [BITMAP_CALL ...]) +``` + **Description:** Xor performs a logical XOR on the results of each `BITMAP_CALL` query passed to it. From dcb411d028705bf92b55d4da9bed07a3d46f6261 Mon Sep 17 00:00:00 2001 From: Matthew Jaffee Date: Fri, 2 Mar 2018 11:00:46 -0600 Subject: [PATCH 5/6] add warnings about sequential ids --- docs/client-libraries.md | 1 + docs/data-model.md | 2 ++ docs/getting-started.md | 5 +++++ 3 files changed, 8 insertions(+) diff --git a/docs/client-libraries.md b/docs/client-libraries.md index 67e36b5dc..a328318ee 100644 --- a/docs/client-libraries.md +++ b/docs/client-libraries.md @@ -10,6 +10,7 @@ nav = [ ## Client Libraries +This section contains example code for client libraries in several languages. Please remember that when modeling your data in Pilosa, it is best to keep row and column ids sequential. It is not wise to use the output of a hash, or randomly distributed ids with Pilosa. ### Go diff --git a/docs/data-model.md b/docs/data-model.md index 5b1741fca..d448a1de8 100644 --- a/docs/data-model.md +++ b/docs/data-model.md @@ -24,6 +24,8 @@ Rows and columns can represent anything (they could even represent the same set Pilosa lays out data first in rows, so queries which get all the set bits in one or many rows, or compute a combining operation on multiple rows such as Intersect or Union are the fastest. Pilosa also has the ability to categorize rows into different "frames" and quickly retrieve the top rows in a frame sorted by the number of bits set in each row. +Please note that Pilosa is most performant when row and column IDs are sequential starting from 0. You can deviate from this to some degree, but if you try to set a bit with column ID 2^63, bad things will start to happen. + ![data model diagram](/img/docs/data-model.svg) ### Index diff --git a/docs/getting-started.md b/docs/getting-started.md index d4802e8ee..f68805f53 100644 --- a/docs/getting-started.md +++ b/docs/getting-started.md @@ -226,6 +226,11 @@ curl localhost:10101/index/repository/query \ {"results":[true]} ``` +Please note that while user ID 99999 may not be sequential with the other column IDs, it is still a relatively low number. +Don't try to use arbitrary 64-bit integers as column or row IDs in Pilosa - this will lead to poor performance, out of memory errors, and more. + + + ### What's Next? You can jump to [Data Model](../data-model/) for an in-depth look at Pilosa's data model, or [Query Language](../query-language/) for more details about **PQL**, the query language of Pilosa. Check out the [Examples](../examples/) page for example implementations of real world use cases for Pilosa. Ready to get going in your favorite language? Have a peek at our small but expanding set of official [Client Libraries](../client-libraries/). From d04530a5c05afc55396adb3bde2e681694c385ff Mon Sep 17 00:00:00 2001 From: Matthew Jaffee Date: Sat, 3 Mar 2018 11:45:48 -0600 Subject: [PATCH 6/6] make container struct description more accurate --- roaring/roaring.go | 14 +++++++++----- 1 file changed, 9 insertions(+), 5 deletions(-) diff --git a/roaring/roaring.go b/roaring/roaring.go index 8ea82af83..3cc537a08 100644 --- a/roaring/roaring.go +++ b/roaring/roaring.go @@ -1002,12 +1002,16 @@ const ArrayMaxSize = 4096 // RunMaxSize represents the maximum size of run length encoded containers. const RunMaxSize = 2048 -// container represents a container for uint32 integers. +// container represents a container for uint16 integers. // -// These are used for storing the low bits. Containers are separated into three -// types depending on cardinality. For containers with less than 4,096 values, -// an array or RLE container is used, depending on the contents. For containers -// with more than 4,096 values, the values are encoded into bitmaps. +// These are used for storing the low bits of numbers in larger sets of uint64. +// The high bits are stored in a container's key which is tracked by a separate +// data structure. Integers in a container can be encoded in one of three ways - +// the encoding used is usually whichever is most compact, though any container +// type should be able to encode any set of integers safely. For containers with +// less than 4,096 values, an array is often used. Containers with long runs of +// integers would use run length encoding, and more random data usually uses +// bitmap encoding. type container struct { mapped bool // mapped directly to a byte slice when true containerType byte // array, bitmap, or run