mirror of
https://github.com/featurebasedb/featurebase.git
synced 2026-08-28 10:54:59 +00:00
* Introduce ServiceManager and Refactor DAX Integration tests The ServiceManager provides an interface with which to manage featurebase (dax) services (mds, queryer, computer). It replaces the confusing interface implementations in /dax/server/server.go (which optionally used pointers to in-process objects to satisfy an interface) with (for now) http implementations. The thought is that even if we're running all services in-process, we should communicate between services over http in order to mirror what we would do in a production environment where the services are running on different nodes. This batch of commits does quit a lot, most of which is captured here: - Added `path` support to `dax.Address`. Address is now a string of the form [scheme]://[host]:[port]/[path]. - Added `Holder.directiveApplied` to determine (in tests) if the computer has completed applying the latest directive. This is somewhat temporary until we improve the mds-to-computer logic. - Removed the "service prefix" code which was prepending client URL paths with the prefix. Instead, the serviceType (mds, queryer, computer[n] is now part of `dax.Address`). - Removed, from the dax config, the top level `StorageMethod` and `StorageDSN` and now just have `MDS.Config.DataDir`. - Added `Computer.Config.N` to specify the number of computers to run in-process. - Moved the `pilosa.MDS` interface to `computer.Registrar`. This is an example of getting the interfaces defined in the right packages. - Added `SnapshotTable()` method to the mds client (to align with its API). - Changed `Balancer.AddJob()` to `Balancer.AddJobs()` to support, for example, adding 256 partitions in a single call. Refactored some of the naive Balancer to account for this. - Added a `Seed` to the top-level config. It's not really useful because of package `crypto/rand`. - Added an in-memory implementation of the DisCo interface and disabled etcd in a computer service. - Create sepearte data-dirs for each in-process computer. - Disabled grpc in dax. - Modified the sql3 test definition format to support multiple insert steps and separate query results (to align with those steps). * Changes necessary to get multiple computer instance running in-process For now the config looks like this: ``` [computer] run = true n = 4 ``` but we can probably just change that to be something like: ``` [computer] run = 4 ``` *Issues found running multiple "computers" in-process* - grpc was trying to bind on the same port - changed GRPCListener from `*net.TCPListener` to `net.Listener` - created a nopListener and set to that for now (i.e. disabled grpc) - etcd was starting more than once - changed dax to use in-memory implementations of the disco interfaces (i.e. stop using etcd) - IDAllocator (which uses boltdb) was trying to open the `idalloc.db` file more than once - realized we have to set separate data-dirs for each holder. that fixed it. * Port dax integration tests to ManagedCommand * Modify Balancer-related methods like AddJob to AddJobs There were (and still are) a lot of places where we were adding on job at a time, even when we had a long list of jobs to add. This resulted in every job add (for example adding 1 of 256 shards) taking ~40ms, or over 10s to create a keyed table. One reason was because each job add was making multiple boltdb transactions. * Port over more dax integration test stuff * Add DirectiveApplied to signify that snapshot/writes have loaded. We use this in tests to avoid using sleeps. This should be considered temporary; we're going to need a more robust solution for determining when a computer node is ready to serve complete data. * Finish porting dax integration tests * Improve godocs * Remove docker-based DAX integration tests. * go mod tidy * Move test/managed.go to avoid package conflicts * Modify IDK integration tests to work with ServiceManager changes This is really just computer -> computer0 And the MDS DataDir config change. * cleanup found during review * echo $CI_COMMIT_REF_SLUG in CI * remove docker image arg, use build instead
95 lines
2.6 KiB
Go
95 lines
2.6 KiB
Go
// Copyright 2021 Molecula Corp. All rights reserved.
|
|
package disco
|
|
|
|
import (
|
|
"context"
|
|
"sort"
|
|
)
|
|
|
|
// Noder is an interface which abstracts the Node slice so that the list of
|
|
// nodes in a cluster can be maintained outside of the cluster struct.
|
|
type Noder interface {
|
|
Nodes() []*Node // Remember: this has to be sorted correctly!!
|
|
PrimaryNodeID(hasher Hasher) string
|
|
|
|
// SetMetadata records the local node's metadata.
|
|
SetMetadata(ctx context.Context, node *Node) error
|
|
|
|
// SetState changes a node to a given state.
|
|
SetState(ctx context.Context, state NodeState) error
|
|
|
|
// ClusterState considers the state of all nodes and gives
|
|
// a general cluster state. The output calculation is as follows:
|
|
// - If any of the nodes are still starting: "STARTING"
|
|
// - If all nodes are up and running: "NORMAL"
|
|
// - If number of DOWN nodes is lower than number of replicas: "DEGRADED"
|
|
// - If number of unresponsive nodes is greater than (or equal to) the number of replicas: "DOWN"
|
|
ClusterState(context.Context) (ClusterState, error)
|
|
}
|
|
|
|
// localNoder is a simple implementation of the Noder interface
|
|
// which maintains an instance of the `nodes` slice.
|
|
type localNoder struct {
|
|
nodes []*Node
|
|
}
|
|
|
|
// NewLocalNoder is a helper function for wrapping an existing slice of Nodes
|
|
// with something which implements Noder.
|
|
func NewLocalNoder(nodes []*Node) *localNoder {
|
|
return &localNoder{
|
|
nodes: nodes,
|
|
}
|
|
}
|
|
|
|
// NewEmptyLocalNoder is an empty Noder used for testing.
|
|
func NewEmptyLocalNoder() *localNoder {
|
|
return &localNoder{}
|
|
}
|
|
|
|
// NewIDNoder is a helper function for wrapping an existing slice of Node IDs
|
|
// with something which implements Noder.
|
|
func NewIDNoder(ids []string) *localNoder {
|
|
nodes := make([]*Node, len(ids))
|
|
for i, id := range ids {
|
|
node := &Node{
|
|
ID: id,
|
|
}
|
|
nodes[i] = node
|
|
}
|
|
|
|
// Nodes must be sorted.
|
|
sort.Sort(ByID(nodes))
|
|
|
|
return &localNoder{
|
|
nodes: nodes,
|
|
}
|
|
}
|
|
|
|
// Nodes implements the Noder interface.
|
|
func (n *localNoder) Nodes() []*Node {
|
|
return n.nodes
|
|
}
|
|
|
|
// PrimaryNodeID implements the Noder interface.
|
|
func (n *localNoder) PrimaryNodeID(hasher Hasher) string {
|
|
snap := NewClusterSnapshot(NewLocalNoder(n.nodes), hasher, "jmp-hash", 1)
|
|
primaryNode := snap.PrimaryFieldTranslationNode()
|
|
if primaryNode == nil {
|
|
return ""
|
|
}
|
|
return primaryNode.ID
|
|
}
|
|
|
|
// ClusterState for localNoder just assumes the cluster is normal.
|
|
func (n *localNoder) ClusterState(context.Context) (ClusterState, error) {
|
|
return ClusterStateNormal, nil
|
|
}
|
|
|
|
func (n *localNoder) SetState(ctx context.Context, state NodeState) error {
|
|
return nil
|
|
}
|
|
|
|
// localNoder doesn't really implement metadata support.
|
|
func (*localNoder) SetMetadata(context.Context, *Node) error {
|
|
return nil
|
|
}
|