Merge pull request #1088 from travisturner/cluster-resize-docs

update docs to include cluster-resize config and instructions
This commit is contained in:
Travis Turner 2018-03-19 15:22:08 -05:00 • committed by GitHub
commit 241150ddeb
No known key found for this signature in database
GPG key ID: 4AEE18F83AFDEB23
3 changed files with 97 additions and 54 deletions

View file

@ -5,6 +5,7 @@ nav = [
"Installing in production",
"Importing and Exporting Data",
"Versioning",
"Resizing the Cluster",
"Backup/restore",
]
+++
@ -93,6 +94,71 @@ The Pilosa server should support PQL versioning using HTTP headers. On each requ
When upgrading, upgrade clients first, followed by server for all Minor and Patch level changes.
### Resizing the Cluster
If you need to increase (or decrease) the capacity of a Pilosa server, you can add or remove nodes to a running cluster at any time. Note that you can only add or remove one node at a time; if you attempt to add multiple nodes at once, those requests will be enqueued and processed serially. Also note that during any resize process, the cluster goes into state `RESIZING` during which all read/write requests are denied. When the cluster returns to state `NORMAL` then read/write operations can resume. The amount of time that the cluster stays in state `RESIZING` depends on the amount of data that needs to be moved during the resize process.
#### Adding a Node
You can add a new, empty node to an existing cluster by starting `pilosa server` on the new node with the correct configuration options. Specifically, you must specify the [cluster coordinator](../configuration/#cluster-coordinator) to be the same as the coordinator on the existing nodes. You must also specify at least one valid [gossip seed](../configuration/#gossip-seeds) (preferably multiple for redundancy). When the new node starts, the coordinator node will receive a `nodeJoin` event indicating that a new node is joining the cluster. At this point, the coordinator will put the cluster into state `RESIZING` and kick off a resize job that instructs all of the nodes in the cluster how to rebalance data to accomodate the additional capacity of the new node. Once the resize job is complete, the coordinator will put the cluster back to state `NORMAL` and ensure that the new node is included in future queries.
If the node is being added to a cluster which contains no data (for example, during startup of a new cluster), the coordinator will bypass the `RESIZING` state and allow the node to join the cluster immediately.
#### Removing a Node
In order to remove a node from a cluster, your cluster must be configured to have a [cluster replicas](../configuration/#cluster-replicas) value of at least 2; if you're removing a node that no longer exists (for example a node that has died), there must be at least one additional replica of the data owned by the dead node in order for the cluster to correctly rebalance itself.
To remove node `localhost:10102` from a cluster having coordinator `localhost:10101`, first determine the ID of the node to be removed. If the node to be removed is still available, you can find the ID by issuing an `/id` request to the node:
``` request
curl localhost:10102/id
```
``` response
40a891fa-243b-4d71-ae24-4f5c78a0f4b1
```
If the node to be removed is no longer available, you can get the IDs of the nodes in the cluster by issuing a `/status` request to any available node:
``` request
curl localhost:10101/status
```
``` response
{
"state":"NORMAL",
"nodes":[
{"id":"24824777-62ec-4151-9fbd-67e4676e317d","uri":{"scheme":"http","host":"localhost","port":10101}}
{"id":"40a891fa-243b-4d71-ae24-4f5c78a0f4b1","uri":{"scheme":"http","host":"localhost","port":10102}}
{"id":"9fab09cc-3c26-4202-9622-d167c84684d9","uri":{"scheme":"http","host":"localhost","port":10103}}
]
}
```
Once you have the ID of the node that you want to remove from the cluster, issue the following request:
```
curl localhost:10101/cluster/resize/remove-node \
-X POST \
-d '{"id": "40a891fa-243b-4d71-ae24-4f5c78a0f4b1"}'
```
At this point, the coordinator will put the cluster into state `RESIZING` and kick off a resize job that instructs all of the nodes in the cluster how to rebalance data to accomodate the reduced capacity of the cluster. Once the resize job is complete, the coordinator will put the cluster back to state `NORMAL` and ensure that the removed node is no longer included in future queries.
Note that you can't directly remove the coordinator node. If you need to remove the coordinator node from the cluster, you must first [make one of the other nodes the coordinator](#changing-the-coordinator).
#### Aborting a Resize Job
If at any point you need to abort an active resize job, you can issue a `POST` request to the `/cluster/resize/abort` endpoint on the coordinator node.
For example, if your coordinator node is `localhost:10101`, then you can run:
```
curl localhost:10101/cluster/resize/abort -X POST
```
This will immediately abort the resize job and return the cluster to state `NORMAL`. Because data is never removed from a node during a resize job (only once a resize job has successfully completed), aborting a resize job will return the cluster back to the state it was in before the resize began.
#### Changing the Coordinator
In order to assign a different node to be the coordinator, you can issue a `/cluster/resize/set-coordinator` request to any node in the cluster. The payload should indicate the ID of the node to be made coordinator.
```
curl localhost:10101/cluster/resize/set-coordinator \
-X POST \
-d '{"id": "9fab09cc-3c26-4202-9622-d167c84684d9"}'
```
### Backup/restore
Pilosa continuously writes out the in-memory bitmap data to disk. This data is organized by Index->Frame->Views->Fragment->numbered slice files. These data files can be routinely backed up to restore nodes in a cluster.

View file

@ -27,23 +27,18 @@ Every command line flag has a corresponding environment variable. The environmen
### Config file
The config file is in the [toml format](https://github.com/toml-lang/toml) and has exactly the same options available as the flags and environment variables. Any flag which contains a dot (".") denotes nesting within the config file, so the two flag `--cluster.replicas=1` looks like this in the config file:
The config file is in the [toml format](https://github.com/toml-lang/toml) and has exactly the same options available as the flags and environment variables. Any flag which contains a dot (".") denotes nesting within the config file, so the two flags `--cluster.coordinator` and `--cluster.replicas=1` look like this in the config file:
```toml
[cluster]
coordinator = true
replicas = 1
```
Any flag that has a value that is a comma separated list on the command line becomes an array in toml. For example `--cluster.hosts=one.pilosa.com:10101,two.pilosa.com:10101` becomes:
```toml
[cluster]
hosts = ["one.pilosa.com:10101", "two.pilosa.com:10101"]
```
### All Options
#### Anti Entropy Interval
* Description: Interval at which the cluster will run its anti-entropy routine which makes sure that all replicas of each fragment are in sync.
* Description: Interval at which the cluster will run its anti-entropy routine which ensures that all replicas of each fragment are in sync.
* Flag: `--anti-entropy.interval="10m0s"`
* Env: `PILOSA_ANTI_ENTROPY_INTERVAL="10m0s"`
* Config:
@ -83,7 +78,7 @@ Any flag that has a value that is a comma separated list on the command line bec
* Config:
```toml
log_path = "/path/to/logfile"
log-path = "/path/to/logfile"
```
#### Max Writes Per Request
@ -132,28 +127,16 @@ Any flag that has a value that is a comma separated list on the command line bec
key = "/var/secret/gossip.key32"
```
#### Cluster Hosts
#### Cluster Coordinator
* Description: List of hosts in the cluster. Multiple hosts should be comma-separated in the flag and env forms.
* Flag: `--cluster.hosts="localhost:10101"`
* Env: `PILOSA_CLUSTER_HOSTS="localhost:10101"`
* Description: Indicates whether the node should act as the coordinator for the cluster. Only one node per cluster should be the coordinator.
* Flag: `cluster.coordinator`
* Env: `PILOSA_CLUSTER_COORDINATOR`
* Config:
```toml
[cluster]
hosts = ["localhost:10101"]
```
#### Cluster Poll Interval
* Description: Polling interval for cluster.
* Flag: `cluster.poll-interval="1m0s"`
* Env: `PILOSA_CLUSTER_POLL_INTERVAL="1m0s"`
* Config:
```toml
[cluster]
poll-interval = "1m0s"
coordinator = true
```
#### Cluster Long Query Time
@ -217,7 +200,8 @@ Any flag that has a value that is a comma separated list on the command line bec
[profile]
cpu-time = "30s"
```
##### Metric Service
#### Metric Service
* Description: Which stats service to use. Choose from [statsd, expvar].
* Flag: `--metric.service=statsd`
* Env: `PILOSA_METRIC_SERVICE=statsd'
@ -228,7 +212,7 @@ Any flag that has a value that is a comma separated list on the command line bec
service = “statsd”
```
##### Metric Host
#### Metric Host
* Description: Address of the StatsD service host.
* Flag: `--metric.host=localhost:8125`
* Env: `PILOSA_METRIC_HOST=localhost:8125'
@ -239,7 +223,7 @@ Any flag that has a value that is a comma separated list on the command line bec
host = "localhost:8125"
```
##### Metric Poll Interval
#### Metric Poll Interval
* Description: Polling interval for runtime metrics.
* Flag: `metric.poll-interval=”0m15s”`
@ -251,7 +235,7 @@ Any flag that has a value that is a comma separated list on the command line bec
poll-interval = "0m15s"
```
##### Metric Diagnostics
#### Metric Diagnostics
* Description: Enable diagnostic reporting. To disable diagnostics set to false.
* Flag: `metric.diagnostics`
@ -264,7 +248,7 @@ Any flag that has a value that is a comma separated list on the command line bec
```
##### TLS Certificate
#### TLS Certificate
* Description: Path to the TLS certificate to use for serving HTTPS. Usually has one of`.crt` or `.pem` extensions.
* Flag: `tls.certificate=/srv/pilosa/certs/server.crt`
@ -276,7 +260,7 @@ Any flag that has a value that is a comma separated list on the command line bec
certificate = "/srv/pilosa/certs/server.crt"
```
##### TLS Certificate Key
#### TLS Certificate Key
* Description: Path to the TLS certificate key to use for serving HTTPS. Usually has the `.key` extension.
* Flag: `tls.key=/srv/pilosa/certs/server.key`
@ -288,7 +272,7 @@ Any flag that has a value that is a comma separated list on the command line bec
key = "/srv/pilosa/certs/server.key"
```
##### TLS Skip Verify
#### TLS Skip Verify
* Description: Disables verification for checking TLS certificates. This configuration item is mainly useful for using self-signed certificates for a Pilosa cluster. Do not use in production since it makes man-in-the-middle attacks trivial.
* Flag: `tls.skip-verify`
@ -315,8 +299,7 @@ A three node cluster running on different hosts could be minimally configured as
[cluster]
replicas = 1
type = "gossip"
hosts = ["node0.pilosa.com:10101","node1.pilosa.com:10101","node2.pilosa.com:10101"]
coordinator = true
#### Node 1
@ -329,8 +312,7 @@ A three node cluster running on different hosts could be minimally configured as
[cluster]
replicas = 1
type = "gossip"
hosts = ["node0.pilosa.com:10101","node1.pilosa.com:10101","node2.pilosa.com:10101"]
coordinator = false
#### Node 2
@ -343,8 +325,7 @@ A three node cluster running on different hosts could be minimally configured as
[cluster]
replicas = 1
type = "gossip"
hosts = ["node0.pilosa.com:10101","node1.pilosa.com:10101","node2.pilosa.com:10101"]
coordinator = false
### Example Cluster Configuration (HTTPS)
@ -363,8 +344,7 @@ The same cluster which uses HTTPS instead of HTTP can be configured as follows.
[cluster]
replicas = 1
type = "gossip"
hosts = ["https://node0.pilosa.com:10101","https://node1.pilosa.com:10101","https://node2.pilosa.com:10101"]
coordinator = true
[tls]
certificate = "/home/pilosa/private/server.crt"
@ -382,8 +362,7 @@ The same cluster which uses HTTPS instead of HTTP can be configured as follows.
[cluster]
replicas = 1
type = "gossip"
hosts = ["https://node0.pilosa.com:10101","https://node1.pilosa.com:10101","https://node2.pilosa.com:10101"]
coordinator = false
[tls]
certificate = "/home/pilosa/private/server.crt"
@ -401,8 +380,7 @@ The same cluster which uses HTTPS instead of HTTP can be configured as follows.
[cluster]
replicas = 1
type = "gossip"
hosts = ["https://node0.pilosa.com:10101","https://node1.pilosa.com:10101","https://node2.pilosa.com:10101"]
coordinator = false
[tls]
certificate = "/home/pilosa/private/server.crt"
@ -424,8 +402,7 @@ You can run a cluster on the same host using the configuration above with a few
[cluster]
replicas = 1
type = "gossip"
hosts = ["https://localhost:10100","https://localhost:10101","https://localhost:10102"]
coordinator = true
[tls]
certificate = "/home/pilosa/private/server.crt"
@ -443,8 +420,7 @@ You can run a cluster on the same host using the configuration above with a few
[cluster]
replicas = 1
type = "gossip"
hosts = ["https://localhost:10100","https://localhost:10101","https://localhost:10102"]
coordinator = false
[tls]
certificate = "/home/pilosa/private/server.crt"
@ -462,8 +438,7 @@ You can run a cluster on the same host using the configuration above with a few
[cluster]
replicas = 1
type = "gossip"
hosts = ["https://localhost:10100","https://localhost:10101","https://localhost:10102"]
coordinator = false
[tls]
certificate = "/home/pilosa/private/server.crt"

View file

@ -35,7 +35,7 @@ Let's make sure Pilosa is running:
curl localhost:10101/status
```
``` response
{"status":{"Nodes":[{"Host":":10101","State":"UP"}]}}
{"state":"NORMAL","nodes":[{"id":"18eb5546-5a1a-4ba4-9c52-b53fbe22317e","uri":{"scheme":"http","host":"localhost","port":10101}}]}
```
### Sample Project
@ -46,7 +46,9 @@ Although Pilosa doesn't keep the data in a tabular format, we still use the term
#### Create the Schema
The queries in this section which are used to set up the indexes in Pilosa just the empty object on success: `{}` - if you would like to verify that a query worked as you expected, you can request the schema as follows:
Note:
The queries in this section which are used to set up the indexes in Pilosa just return the empty object on success: `{}` - if you would like to verify that a query worked as you expected, you can request the schema as follows:
``` request
curl localhost:10101/schema
```
@ -107,7 +109,7 @@ docker cp language.csv pilosa:/language.csv
docker exec -it pilosa /pilosa import -i repository -f language /language.csv
```
Note that, both the user IDs and the repository IDs were remapped to sequential integers in the data files, they don't correspond to actual Github IDs anymore. You can check out `languages.txt` to see the mapping for languages.
Note that both the user IDs and the repository IDs were remapped to sequential integers in the data files, they don't correspond to actual Github IDs anymore. You can check out [languages.txt](https://github.com/pilosa/getting-started/blob/master/languages.txt) to see the mapping for languages.
### Input Definition
Alternatively Pilosa can import JSON data using an [Input Definition](../input-definition/) describing the schema and ETL rules to process the data.